AI video is Brownian motion played backwards.
D9 · Image & Video Generation
Generation is the reverse of destruction. Add noise to an image in small steps until nothing is left, train a model to undo one step, and you can start from pure noise and walk back to a picture. This chapter builds that idea from autoencoders through VAEs and GANs to diffusion, so you see what each one fixed. You will also learn why guidance changes the look of an output, why latent diffusion made all of this affordable, what makes a video hold together, and what a published score cannot see.
- Squeeze, Then Rebuild
- Walking Through Latent Space
- Variational Autoencoders
- Two Networks in a Fight
- Why GANs Collapse
- Destroying an Image on Purpose
- Learning to Undo One Step
- Following the Data Uphill
- Sampling Schedules
- Conditioning and Guidance
- Diffusion in Latent Space
- From Words to Pictures
- Making Time Consistent
- Scoring a Picture Maker
- What These Models Understand