December 13, 2025


Mountain and Tree Flip Image
Random seed: 1024
Prompts:
Reflection: More inference steps improved quality; long/abstract prompts introduced more mistakes/strangeness.








Increasing noise on the Campanile via the forward diffusion process.



Gaussian blur struggles to recover structure from heavily noised inputs.



Single UNet denoise pass recovers detail but degrades as noise increases.






Strided timesteps progressively denoise toward the clean estimate.









Unconditional samples from noise (prompt: "a high quality photo").





CFG boosts fidelity; sampling with guidance scale > 1.





Denoising with different start noise levels (higher = less edits).





















Progression from low to high noise starts





















Filling masked regions while preserving context. From left to right: Original, Mask, Hole, Inpainted.












Guided edits from different noise levels; Lower noise looks more like text, higher noise looks more lke image.
Prompt: "A pixel art style image, with high pixel density"





















Images that flip into a different concept when rotated 180°. Each row shows the original orientation and flipped version at different resolutions.
64×64 resolution



256×256 resolution






Combining low- and high-frequency noise estimates to blend two prompts.

Palm × Map hybrid

Balloon × Pig hybrid

Landscape × Bookshelf hybrid
Implementing and training UNet-based denoising and flow matching models on MNIST.
Visualizing the noising process and training a UNet to denoise MNIST digits at σ = 0.5.

Noising process at different σ levels

Epoch 1

Epoch 5

Training loss curve
Testing the denoiser on varying noise levels it wasn't trained on.

Performance across different σ values
Training a denoiser to denoise pure noise. The patterns observed in the generated outputs seem to resemble an average of the MNIST digits stacked on top of each other This is likely because there is no prior information about the digits in each noise sample. Thus, to minimize loss, the denoiser has to learn the average of the digits.

Epoch 1

Epoch 5

Training loss over 5 epochs
Training a time-conditioned UNet to predict flow from noisy to clean data.

Epoch 1

Epoch 5

Epoch 10

Training loss curve
Adding class conditioning to control digit generation with classifier-free guidance.

Epoch 1

Epoch 5

Epoch 10

Training loss curve
Compensating for removal of exponential learning rate decay by lowering the learning rate to 5e-3.

Epoch 1

Epoch 5

Epoch 10

Training loss without scheduler