←Back to CS 180
P5

Project 5: Diffusion Models

December 13, 2025

Diffusion ModelsNeural NetworksGenerative Models

Overview

Mountain and Tree Flip ImageMountain and Tree Flip Image

Mountain and Tree Flip Image

Part 0: Text-to-Image (DeepFloyd)

Random seed: 1024

Prompts:

  • hands using a figure skate as a knife to cut into a hyper-realistic raw steak-shaped cake
  • an overhead drone-scan style view of an abandoned ice-skating rink at dawn, the ice cracked and reflecting the first light, faint figure skating loops etched on it, muted pastel palette
  • a translucent hawk soaring over mountains, its wings filled with floating code snippets instead of feathers
  • a mixed-media style portrait of a girl running statistical regressions on a foggy mountain trail, numbers dissolving into mist

Reflection: More inference steps improved quality; long/abstract prompts introduced more mistakes/strangeness.

Steak prompt 30 stepsSteak prompt 80 steps
Steak cake (30, 80 steps)
Rink prompt 50 stepsRink prompt 100 steps
Abandoned rink (50, 100 steps)
Hawk prompt 50 stepsHawk prompt 100 steps
Hawk (50, 100 steps)
Runner prompt 20 stepsRunner prompt 50 steps
Runner (20, 50 steps)

1.1 Forward Process (Adding Noise)

Increasing noise on the Campanile via the forward diffusion process.

Campanile noisy t=250
t = 250
Campanile noisy t=500
t = 500
Campanile noisy t=750
t = 750

Gaussian blur struggles to recover structure from heavily noised inputs.

Gaussian blur t=250
t = 250
Gaussian blur t=500
t = 500
Gaussian blur t=750
t = 750

1.3 One-Step Denoising

Single UNet denoise pass recovers detail but degrades as noise increases.

Campanile noisy t=250
t = 250
Campanile noisy t=500
t = 500
Campanile noisy t=750
t = 750
One-step denoise t=250
t = 250
One-step denoise t=500
t = 500
One-step denoise t=750
t = 750

1.4 Iterative Denoising

Strided timesteps progressively denoise toward the clean estimate.

Iterative step 90
t=90
Iterative step 240
t=240
Iterative step 390
t=390
Iterative step 540
t=540
Iterative step 690
t=690
Campanile original
Original
One-step at heavy noise
One-step (t≈990)
Gaussian blur
Gaussian blur
Iterative final
Iterative final

1.5 Diffusion Sampling

Unconditional samples from noise (prompt: "a high quality photo").

Sample 1
1
Sample 2
2
Sample 3
3
Sample 4
4
Sample 5
5

1.6 Classifier-Free Guidance

CFG boosts fidelity; sampling with guidance scale > 1.

CFG sample 1
1
CFG sample 2
2
CFG sample 3
3
CFG sample 4
4
CFG sample 5
5

1.7 SDEdit / Image-to-Image

Denoising with different start noise levels (higher = less edits).

Campanile i_start=1
i_start=1
Campanile i_start=3
i_start=3
Campanile i_start=5
i_start=5
Campanile i_start=7
i_start=7
Campanile i_start=10
i_start=10
Campanile i_start=20
i_start=20
Campanile original
Original
Dino i_start=1
i_start=1
Dino i_start=3
i_start=3
Dino i_start=5
i_start=5
Dino i_start=7
i_start=7
Dino i_start=10
i_start=10
Dino i_start=20
i_start=20
Dino original
Original
Fox i_start=1
i_start=1
Fox i_start=3
i_start=3
Fox i_start=5
i_start=5
Fox i_start=7
i_start=7
Fox i_start=10
i_start=10
Fox i_start=20
i_start=20
Fox original
Original

1.7.1 Editing Web & Hand-Drawn Inputs

Progression from low to high noise starts

Cat-fish i_start=1
1
Cat-fish i_start=3
3
Cat-fish i_start=5
5
Cat-fish i_start=7
7
Cat-fish i_start=10
10
Cat-fish i_start=20
20
Original cat
Original
Bear i_start=1
1
Bear i_start=3
3
Bear i_start=5
5
Bear i_start=7
7
Bear i_start=10
10
Bear i_start=20
20
Original bear
Original
Hand sketch i_start=1
1
Hand sketch i_start=3
3
Hand sketch i_start=5
5
Hand sketch i_start=7
7
Hand sketch i_start=10
10
Hand sketch i_start=20
20
Original
Original

1.7.2 Inpainting

Filling masked regions while preserving context. From left to right: Original, Mask, Hole, Inpainted.

Original Campanile
Original
Campanile mask
Mask
Campanile hole
Hole
Campanile inpainted
Inpainted
Original Elmo
Original
Elmo mask
Mask
Elmo hole
Hole
Elmo inpainted
Inpainted
Original Ocean
Original
Ocean mask
Mask
Ocean hole
Hole
Ocean inpainted
Inpainted

1.7.3 Text-Conditional Translation

Guided edits from different noise levels; Lower noise looks more like text, higher noise looks more lke image.
Prompt: "A pixel art style image, with high pixel density"

Cat text edit 1
1
Cat text edit 3
3
Cat text edit 5
5
Cat text edit 7
7
Cat text edit 10
10
Cat text edit 20
20
Original cat
Original
Cow text edit 1
1
Cow text edit 3
3
Cow text edit 5
5
Cow text edit 7
7
Cow text edit 10
10
Cow text edit 20
20
Original cow
Original
Campanile text edit 1
1
Campanile text edit 3
3
Campanile text edit 5
5
Campanile text edit 7
7
Campanile text edit 10
10
Campanile text edit 20
20
Original campanile
Original

1.8 Visual Anagrams

Images that flip into a different concept when rotated 180°. Each row shows the original orientation and flipped version at different resolutions.

64×64 resolution

Man and campfire flip
Old man ↔ Campfire
Clock and face flip
Clocks ↔ face
Mountain and tree flip
Mountain ↔ Oak tree

256×256 resolution

Old man
An oil painting of an old man
Campfire
An oil painting of people around a campfire
Melted clocks
An oil painting of a pile of melted clocks
Crying face
An oil painting of a surreal crying face
Mountain
A water color drawing of a mountain
Oak tree
A water color drawing of an oak tree

1.9 Hybrid Images

Combining low- and high-frequency noise estimates to blend two prompts.

Hybrid palm + map

Palm × Map hybrid

Hybrid balloon + pig

Balloon × Pig hybrid

Hybrid landscape + bookshelf

Landscape × Bookshelf hybrid

Part B: Training a Diffusion Model from Scratch

Implementing and training UNet-based denoising and flow matching models on MNIST.

Part 1.2: Training a UNet-based Denoiser

Visualizing the noising process and training a UNet to denoise MNIST digits at σ = 0.5.

Noising process visualization

Noising process at different σ levels

Uncond epoch 1

Epoch 1

Uncond epoch 5

Epoch 5

Uncond training loss

Training loss curve

Part 1.2.2: Out-of-Distribution Testing

Testing the denoiser on varying noise levels it wasn't trained on.

Out-of-distribution results

Performance across different σ values

Part 1.2.3: Denoising Pure Noise

Training a denoiser to denoise pure noise. The patterns observed in the generated outputs seem to resemble an average of the MNIST digits stacked on top of each other This is likely because there is no prior information about the digits in each noise sample. Thus, to minimize loss, the denoiser has to learn the average of the digits.

Results after epoch 1

Epoch 1

Results after epoch 5

Epoch 5

Training loss curve

Training loss over 5 epochs

Part 2: Flow Matching Model

2.3: Time-Conditioned UNet

Training a time-conditioned UNet to predict flow from noisy to clean data.

Epoch 1 samples

Epoch 1

Epoch 5 samples

Epoch 5

Epoch 10 samples

Epoch 10

Time-conditioned training loss

Training loss curve

2.6: Class-Conditioned UNet with CFG

Adding class conditioning to control digit generation with classifier-free guidance.

Class-conditioned epoch 1

Epoch 1

Class-conditioned epoch 5

Epoch 5

Class-conditioned epoch 10

Epoch 10

Class-conditioned training loss

Training loss curve

Training Without LR Scheduler

Compensating for removal of exponential learning rate decay by lowering the learning rate to 5e-3.

No decay epoch 1

Epoch 1

No decay epoch 5

Epoch 5

No decay epoch 10

Epoch 10

No decay training loss

Training loss without scheduler