September 26, 2025

A zebra car I saw on my run in spatial domain

...vs frequency domain
Understanding images from a frequency perspective rather than spatial perspective is crucial in digital image processing. Images are made up of different frequency components, with low frequencies capturing the general structure of an image and high frequencies capturing small/fine details. By manipulating the constituent frequencies of images, we can blur, sharpen, and blend images seamlessly.
The goal of this assignment is to demonstrate how image processing in the frequency domain can be used to achieve interesting visual effects.
Mathematically, convolution is defined as:
Convolution is an associative and communative operation, i.e.
The following are a couple of filters/kernels:

Original Image (f): my sister and I skating together again!

f∗Dx

f∗Dy

f∗Box
* Should look blurry, but here, it is still relatively clear since the kernel (3x3) is small compared to my image
In practice, convolution is a very commonly used operation, leading to many performance optimizations (through formulation of convolution as matrix-multiplication followed by parallelization) Let's compare the runtime it takes to convolve an image and filter using SciPy's convolve2d, a function from a prebuild library, with two naive implementations.
Scipy convolution time: 0.1677 seconds
VERSUS
def convolve_four_loops(image, kernel):
... (handling of padding) ...
# In my implementation, I used zero-padding, i.e. just add zeros to the edge of the original image
# But I found that when using Sci-py's convolve2d, 'symm' padding, i.e. mirror the edge pixels,
# was the best for the processing I was doing
for i in range(ih):
for j in range(iw):
sum = 0
for m in range(kh):
for n in range(kw):
sum += kernel[m, n] * padded_image[i + m, j + n] # this is scalar multiplication
output[i, j] = sum
return outputFour loops convolution time: 19.7630 secondsdef convolve_two_loops(image, kernel):
... (handling of padding) ...
for i in range(ih):
for j in range(iw):
region = padded_image[i:i+kh, j:j+kw]
output[i, j] = np.sum(region * kernel) # this is element-wise multiplication followed by sum of all elements
return outputTwo loops convolution time: 19.6246 secondsThe finite difference operator is a simple filter that approximates the derivative of an image in a specific direction. i.e.

Original Image (f): "camera_man.png"

f∗Dx

f∗Dy

∥∇f∥2

∥∇f∥2≥0.32
(thresholded gradient magnitude)
* Note how the thresholded gradient magnitude image clearly highlights the strong edges in the image.
In real-world images, noise can create false edges and make it difficult to accurately detect true edges. To mitigate this, we can smooth the image using a Gaussian filter before applying the derivative operator. This combination (of Gaussian smoothing and finite difference operator) is known as the Derivative of Gaussian (DoG) filter.

fsmoothed∗Dx

fsmoothed∗Dy

∥∇fsmoothed∥2

∥∇fsmoothed∥2≥0.1
(thresholded gradient magnitude)
Comparing these derivatives and magnitudes with those from a plain Finite Difference Operator convolution, we can see how DoG creates a more robust edge detection filter.
To understand why the Derivative of Gaussian (DoG) filter reduces high-frequency noise while preserving important low-frequency structures in the image (thus making edges more distinct), let's break down the process into two steps:
* I chose σ=2.0 and a k=int(2∗np.ceil(3∗std)+1) (to capture 99% of the Gaussian distribution) for the above results because it seemed to get rid of enough noise while preserving the low-frequency edges of the camera-man.
** To choose the thresholds for the gradient magnitude, I tried a few values and chose the ones that seem to throw away all the background (e.g. grass) and keep the edges of the camera-man.
The Gaussian filter we used earlier acts as low pass filter that retains only the low frequencies. If we subtract the blurred version (only low frequencies) from the original image, we get the high frequencies of the image. An image often looks "sharper" if it has more high frequencies, so we can make images look "sharper" by adding these high frequencies to the original image.
Let's visualize the process of the unsharp filter (from left to right):

Original Image (f): "taj_mahal.jpg"

f∗Gσ=2
(Low Frequencies)

f−(f∗Gσ=2)
(High Frequencies)

f+α(f−(f∗Gσ=2))
(Sharpened with α=1)
We can also combine this process into a single convolution operation called the unsharp mask filter.
Let f denote a grayscale 2D image, alpha denote the sharpening parameter, Gsigma denote a Gaussian filter, and delta denote a 2-D unit impulse. Starting from the definition of the unsharp procedure, we apply the properties of convolution to get a single convolution operation:
alpha denotes the sharpening parameter, i.e. the "strength" of sharpening. The larger the alpha, the more of the high-frequency components we are adding, creating a stronger sharpening effect.

Guilin, China

alpha=1,sigma=2

alpha=4,sigma=2

alpha=10,sigma=2

Peach from Oregon

alpha=1,sigma=2

alpha=4,sigma=2

alpha=10,sigma=2

Ducks and geese in Dongshan, China

Sharpened (alpha=2,sigma=2)

Sharpened → Blurred

Sharpened → Blurred → Sharpened
* Looks less sharp compared to original and to first sharpening

Nutmeg

Derek

Nutmeg and Derek Hybrid (BW-BW)
σhigh=5,σlow=9,α=1

Nara, Japan Deer

Tilden, Berkeley Cow

Nara, Japan Deer and Tilden, Berkeley Cow Hybrid (BW-BW)
σhigh=5,σlow=10,α=1

Carlos (my cat)

Self-portrait

Carlos and Iana Self-Portrait Hybrid (BW-BW)
σhigh=5,σlow=10,α=1

Carlos (my cat)

Self-portrait

Carlos and Iana Self-Portrait Hybrid (BW-BW)
σhigh=5,σlow=10,α=1

Carlos Aligned (c)

Self-portrait Aligned (s)

High Frequency Component
c−(c∗Gσhigh=10)

Low Frequency Component
s∗Gσlow=5
I found that the output hybrid image seems to look better when using color only for the high-frequency component, and grayscale for the low-frequency component

Nutmeg and Derek Hybrid (RGB-BW)
σhigh=5,σlow=9,α=0.5

Nara, Japan Deer and Tilden, Berkeley Cow Hybrid (RGB-BW)
σhigh=5,σlow=10,α=0.5

Carlos and Iana Self-Portrait Hybrid (RGB-BW)
σhigh=5,σlow=10,α=0.5

Recreation of Figure 3.42 in Szelski

Oraple (vertical seam mask)

Shanghai -- Day and Night (horizontal seam mask)

Akaza -- Before and After (circular mask)