←Back to CS 180
P2

Project 2: Fun with Filters and Frequencies

September 26, 2025

Image Frequencies

Overview

A zebra car I saw on my run in spatial domain

A zebra car I saw on my run in spatial domain

...vs frequency domain

...vs frequency domain

Understanding images from a frequency perspective rather than spatial perspective is crucial in digital image processing. Images are made up of different frequency components, with low frequencies capturing the general structure of an image and high frequencies capturing small/fine details. By manipulating the constituent frequencies of images, we can blur, sharpen, and blend images seamlessly.

The goal of this assignment is to demonstrate how image processing in the frequency domain can be used to achieve interesting visual effects.

Convolutions

1.1 From Scratch!

Mathematically, convolution is defined as:

(A∗B)[i,j]=∑l=−∞∞∑l=−∞∞A[k,l]B[i−k,j−l] (A * B) [i, j]= \sum\limits_{l=-\infty}^{\infty} \sum\limits_{l=-\infty}^\infty A [k, l] B[i - k, j - l]
for convolution of two single-channel images/filters, A&BA \hspace{0.1cm}\& \hspace{0.1cm}B, both of dimension (H,W)(H, W)

Convolution is an associative and communative operation, i.e.

A∗B=B∗A&A∗(B−C)=A∗B−A∗C A * B = B * A \\ \& \\ A * (B - C) = A * B - A * C

Let's take a look at what convolution does!

The following are a couple of filters/kernels:

Dx=[10−1]Dy=[10−1]Box=19[111111111] \begin{align*} D_x &= \begin{bmatrix} 1 & 0 & -1 \end{bmatrix} \hspace{0.5cm} D_y &= \begin{bmatrix} 1 \\ 0 \\ -1 \end{bmatrix} \hspace{0.5cm} Box &= \frac{1}{9}\begin{bmatrix} 1 & 1 & 1 \\ 1 & 1 & 1 \\ 1 & 1 & 1 \\ \end{bmatrix} \end{align*}
And an image convolved with these filters/kernels:Original Image (f): my sister and I skating together again!

Original Image (ff): my sister and I skating together again!

f * D_x

f∗Dxf * D_x

f * D_

f∗Dyf * D_y

f * Box

f∗Boxf * Box
* Should look blurry, but here, it is still relatively clear since the kernel (3x3) is small compared to my image

Code Implementation

In practice, convolution is a very commonly used operation, leading to many performance optimizations (through formulation of convolution as matrix-multiplication followed by parallelization) Let's compare the runtime it takes to convolve an image and filter using SciPy's convolve2d, a function from a prebuild library, with two naive implementations.


Scipy convolution time: 0.1677 seconds

VERSUS

def convolve_four_loops(image, kernel):

  ... (handling of padding) ...
  # In my implementation, I used zero-padding, i.e. just add zeros to the edge of the original image
  # But I found that when using Sci-py's convolve2d, 'symm' padding, i.e. mirror the edge pixels,
  # was the best for the processing I was doing

  for i in range(ih):
      for j in range(iw):
          sum = 0
          for m in range(kh):
              for n in range(kw):
                  sum += kernel[m, n] * padded_image[i + m, j + n] # this is scalar multiplication
          output[i, j] = sum

    return output
Four loops convolution time: 19.7630 seconds
def convolve_two_loops(image, kernel):
              
  ... (handling of padding) ...

    for i in range(ih):
        for j in range(iw):
            region = padded_image[i:i+kh, j:j+kw]
            output[i, j] = np.sum(region * kernel) # this is element-wise multiplication followed by sum of all elements


    return output
Two loops convolution time: 19.6246 seconds

Filters!

1.2 Finite Difference Operator

The finite difference operator is a simple filter that approximates the derivative of an image in a specific direction. i.e.

∂f∂x≈f∗Dx∂f∂y≈f∗Dx \begin{align*} \frac{\partial f}{\partial x} \approx f * D_x \hspace{0.5cm} \frac{\partial f}{\partial y} \approx f * D_x \end{align*}
Each operator highlights regions of rapid intensity change (in the x and y directions, respectively), which often correspond to edges in the image. To detect edges regardless of their orientation, we can compute the derivative magnitude:
∥∇f∥2=(∂f∂x)2+(∂f∂y)2\Vert \nabla f \Vert_2 = \sqrt{\left(\frac{\partial f}{\partial x}\right)^2 + \left(\frac{\partial f}{\partial y}\right)^2}
Let's visualize these derivatives and magnitudes for f=f = "camera_man.png", as well as a thresholded version of the gradient magnitude

Original Image f =  'camera_man.png'

Original Image (ff): "camera_man.png"

f * D_x

f∗Dxf * D_x

f * D_y

f∗Dyf * D_y

\Vert \nabla f \Vert_2

∥∇f∥2\Vert \nabla f \Vert_2

\Vert \nabla f \Vert_2 \geq 0.32

∥∇f∥2≥0.32\Vert \nabla f \Vert_2 \geq 0.32
(thresholded gradient magnitude)

* Note how the thresholded gradient magnitude image clearly highlights the strong edges in the image.

1.3 Derivative of Gaussian (DoG) Filters

In real-world images, noise can create false edges and make it difficult to accurately detect true edges. To mitigate this, we can smooth the image using a Gaussian filter before applying the derivative operator. This combination (of Gaussian smoothing and finite difference operator) is known as the Derivative of Gaussian (DoG) filter.

f_{smoothed} * D_x

fsmoothed∗Dxf_{smoothed} * D_x

f_{smoothed} * D_y

fsmoothed∗Dyf_{smoothed} * D_y

\Vert \nabla f_{smoothed} \Vert_2

∥∇fsmoothed∥2\Vert \nabla f_{smoothed} \Vert_2

\Vert \nabla f_{smoothed} \Vert_2 \geq 0.1

∥∇fsmoothed∥2≥0.1\Vert \nabla f_{smoothed} \Vert_2 \geq 0.1
(thresholded gradient magnitude)

Comparing these derivatives and magnitudes with those from a plain Finite Difference Operator convolution, we can see how DoG creates a more robust edge detection filter.

To understand why the Derivative of Gaussian (DoG) filter reduces high-frequency noise while preserving important low-frequency structures in the image (thus making edges more distinct), let's break down the process into two steps:

  1. smoothing the image to reduce noise that obscures or distracts from edges (remove high-frequencies)
    fsmoothed=f∗Gσ \begin{align*} f_{smoothed} &= f * G_{\sigma} \\ \end{align*}
  2. applying a derivative operation on the smoothed image to detect edges.
    ∂∂x(f∗Gσ)≈fsmoothed∗Dx=(f∗Gσ)∗Dx∂∂y(f∗Gσ)≈fsmoothed∗Dy=(f∗Gσ)∗Dy \begin{align*} \frac{\partial}{\partial x} (f * G_{\sigma}) &\approx f_{smoothed} * D_x = (f * G_{\sigma}) * D_x\\ \frac{\partial}{\partial y} (f * G_{\sigma}) &\approx f_{smoothed} * D_y = (f * G_{\sigma}) * D_y\\ \end{align*}
Using the associative property of convolution, we can rewrite the above as:
∂∂x(f∗Gσ)≈f∗(Gσ∗Dx)=f∗DoGx∂∂y(f∗Gσ)≈f∗(Gσ∗Dy)=f∗DoGy \begin{align*} \frac{\partial}{\partial x} (f * G_{\sigma}) &\approx f * (G_{\sigma} * D_x) = f * DoG_x \\ \frac{\partial}{\partial y} (f * G_{\sigma}) &\approx f * (G_{\sigma} * D_y) = f * DoG_y \\ \end{align*}
This means that we can first convolve the Gaussian filter with the derivative filter to create a single DoG filter, and then convolve the image with this DoG filter to get the same result. This is computationally more efficient, especially for larger images and filters. To verify that both methods yield the same result, I computed the maximum difference between the two methods' results for both x and y derivatives, which were both very small:

Max difference in X derivatives: 0.004088
Max difference in Y derivatives: 0.004337

* I chose σ=2.0\sigma = 2.0 and a k=int(2∗np.ceil(3∗std)+1)k = int(2 * np.ceil(3 * std) + 1) (to capture 99% of the Gaussian distribution) for the above results because it seemed to get rid of enough noise while preserving the low-frequency edges of the camera-man.
** To choose the thresholds for the gradient magnitude, I tried a few values and chose the ones that seem to throw away all the background (e.g. grass) and keep the edges of the camera-man.

Image "Sharpening"

2.1 How it works

The Gaussian filter we used earlier acts as low pass filter that retains only the low frequencies. If we subtract the blurred version (only low frequencies) from the original image, we get the high frequencies of the image. An image often looks "sharper" if it has more high frequencies, so we can make images look "sharper" by adding these high frequencies to the original image.

Let's visualize the process of the unsharp filter (from left to right):

Original Image (f):'taj_mahal.jpg'

Original Image (ff): "taj_mahal.jpg"

Low Frequencies

f∗Gσ=2f * G_{\sigma = 2}
(Low Frequencies)

High Frequencies

f−(f∗Gσ=2)f - (f * G_{\sigma = 2})
(High Frequencies)

Sharpened Image

f+α(f−(f∗Gσ=2))f + \alpha(f - (f * G_{\sigma = 2}))
(Sharpened with α=1\alpha=1)

Unsharp Mask Filter

We can also combine this process into a single convolution operation called the unsharp mask filter.

Let ff denote a grayscale 2D image, alphaalpha denote the sharpening parameter, GsigmaG_{sigma} denote a Gaussian filter, and deltadelta denote a 2-D unit impulse. Starting from the definition of the unsharp procedure, we apply the properties of convolution to get a single convolution operation:

f+α(f−f∗g)=f∗δ+α(f∗δ−f∗g)=f∗δ+α(f∗δ)−α(f∗g)=f∗δ+f∗(αδ)−f∗(αg)=f∗((1+α)δ−αg)\begin{align} f + \alpha(f - f * g) &= f * \delta + \alpha(f * \delta - f * g) \\ &= f * \delta + \alpha(f * \delta) - \alpha(f * g) \\ &= f * \delta + f * (\alpha\delta) - f * (\alpha g) \\ &= f * ((1 + \alpha)\delta - \alpha g) \end{align}

Visualizing the effects of the sharpening parameter, alphaalpha

alphaalpha denotes the sharpening parameter, i.e. the "strength" of sharpening. The larger the alphaalpha, the more of the high-frequency components we are adding, creating a stronger sharpening effect.

Guilin, China

Guilin, China

Guilin, China α=1

alpha=1,sigma=2alpha = 1 , sigma = 2

Guilin, China α=4

alpha=4,sigma=2alpha = 4 , sigma = 2

Guilin, China α=10

alpha=10,sigma=2alpha = 10 , sigma = 2

Peach from Oregon

Peach from Oregon

Peach α=1

alpha=1,sigma=2alpha = 1 , sigma = 2

Peach α=4

alpha=4,sigma=2alpha = 4 , sigma = 2

Peach α=10

alpha=10,sigma=2alpha = 10 , sigma = 2

Experiment: Sharpen an Image, Blur It, then Sharpen It Again

Ducks and geese in Dongshan, China

Ducks and geese in Dongshan, China

Sharpened

Sharpened (alpha=2,sigma=2alpha = 2 , sigma = 2)

Sharpened → Blurred

Sharpened → Blurred

Sharpened → Blurred → Sharpened

Sharpened → Blurred → Sharpened
* Looks less sharp compared to original and to first sharpening

Hybrid Images

2.2

Nutmeg

Nutmeg

Derek

Derek

Nutmeg and Derek Hybrid (BW-BW)

Nutmeg and Derek Hybrid (BW-BW)
σhigh=5,σlow=9,α=1\sigma_{high} = 5, \sigma_{low} = 9, \alpha = 1

Nara, Japan Deer

Nara, Japan Deer

Tilden, Berkeley Cow

Tilden, Berkeley Cow

Nara, Japan Deer and Tilden, Berkeley Cow Hybrid (BW-BW)

Nara, Japan Deer and Tilden, Berkeley Cow Hybrid (BW-BW)
σhigh=5,σlow=10,α=1\sigma_{high} = 5, \sigma_{low} = 10, \alpha = 1

Carlos (my cat)

Carlos (my cat)

Self-portrait

Self-portrait

Carlos and Iana Self-Portrait Hybrid (BW-BW)

Carlos and Iana Self-Portrait Hybrid (BW-BW)
σhigh=5,σlow=10,α=1\sigma_{high} = 5, \sigma_{low} = 10, \alpha = 1

Carlos (my cat)

Carlos (my cat)

Self-portrait

Self-portrait

Carlos and Iana Self-Portrait Hybrid (BW-BW)

Carlos and Iana Self-Portrait Hybrid (BW-BW)
σhigh=5,σlow=10,α=1\sigma_{high} = 5, \sigma_{low} = 10, \alpha = 1

Carlos Aligned

Carlos Aligned (cc)

Self-portrait Aligned

Self-portrait Aligned (ss)

High Frequency Component
c−(c∗Gσhigh=10)c - (c * G_{\sigma_{high} = 10})

Self-portrait

Low Frequency Component
s∗Gσlow=5s * G_{\sigma_{low} = 5}

I found that the output hybrid image seems to look better when using color only for the high-frequency component, and grayscale for the low-frequency component

Nutmeg and Derek Hybrid (RGB-BW)

Nutmeg and Derek Hybrid (RGB-BW)
σhigh=5,σlow=9,α=0.5\sigma_{high} = 5, \sigma_{low} = 9, \alpha = 0.5

Nara, Japan Deer and Tilden, Berkeley Cow Hybrid (RGB-BW)

Nara, Japan Deer and Tilden, Berkeley Cow Hybrid (RGB-BW)
σhigh=5,σlow=10,α=0.5\sigma_{high} = 5, \sigma_{low} = 10, \alpha = 0.5

Carlos and Iana Self-Portrait Hybrid (BW-BW)

Carlos and Iana Self-Portrait Hybrid (RGB-BW)
σhigh=5,σlow=10,α=0.5\sigma_{high} = 5, \sigma_{low} = 10, \alpha = 0.5

Gaussian and Laplacian Stacks

2.3

Recreation of Figure 3.42 in Szelski

Recreation of Figure 3.42 in Szelski

2.4 Multiresolution Blending (a.k.a. the oraple!)

Oraple

Oraple (vertical seam mask)

Shanghai -- Day and Night

Shanghai -- Day and Night (horizontal seam mask)

Akaza -- Before and After

Akaza -- Before and After (circular mask)