Stage 3 · Expert · D9
Image & Video Generation
Most AI pictures and videos start as pure static.
15 lessons · 178 minDemanding
About this chapter
Generation is the reverse of destruction. Add noise to an image in small steps until nothing is left, train a model to undo one step, and you can start from pure noise and walk back to a picture. This chapter builds that idea from autoencoders through VAEs and GANs to diffusion, so you see what each one fixed. You will also learn why guidance changes the look of an output, why latent diffusion made all of this affordable, what makes a video hold together, and what a published score cannot see.
What you will be able to do
- 110 min
Squeeze, Then Rebuild
Train an autoencoder and read what the bottleneck kept.
- 211 min
Walking Through Latent Space
Interpolate between two points in latent space and describe the path.
- 312 min
Variational Autoencoders
Explain why a VAE predicts a distribution instead of a point.
- 412 min
Two Networks in a Fight
Describe the generator and discriminator game and its equilibrium.
- 512 min
Why GANs Collapse
Recognize mode collapse and name two common mitigations.
- 612 min
Destroying an Image on Purpose
Add noise over a schedule and describe the image at each stage.
- 713 min
Learning to Undo One Step
Explain what the network predicts at each denoising step.
- 812 min
Following the Data Uphill
Read the score as a direction toward more probable images.
- 911 min
Sampling Schedules
Trade steps against quality and pick a sampler for a budget.
- 1013 min
Conditioning and Guidance
Explain classifier-free guidance and predict what a high scale does.
- 1112 min
Diffusion in Latent Space
Explain why running diffusion on latents cut the cost so sharply.
- 1212 min
From Words to Pictures
Trace how a text encoder steers image generation, and where it misreads.
- 1312 min
Making Time Consistent
Explain what changes when frames must agree with each other.
- 1412 min
Scoring a Picture Maker
Read an FID number honestly and name three things it cannot see.
- 1512 min
What These Models Understand
Judge from failure cases what the model represents and what it fakes.