Skip to content

Stage 3 · Expert · D9

Image & Video Generation

Most AI pictures and videos start as pure static.

15 lessons · 178 minDemanding

About this chapter

Generation is the reverse of destruction. Add noise to an image in small steps until nothing is left, train a model to undo one step, and you can start from pure noise and walk back to a picture. This chapter builds that idea from autoencoders through VAEs and GANs to diffusion, so you see what each one fixed. You will also learn why guidance changes the look of an output, why latent diffusion made all of this affordable, what makes a video hold together, and what a published score cannot see.

What you will be able to do

  1. 1

    Squeeze, Then Rebuild

    Train an autoencoder and read what the bottleneck kept.

    10 min
  2. 2

    Walking Through Latent Space

    Interpolate between two points in latent space and describe the path.

    11 min
  3. 3

    Variational Autoencoders

    Explain why a VAE predicts a distribution instead of a point.

    12 min
  4. 4

    Two Networks in a Fight

    Describe the generator and discriminator game and its equilibrium.

    12 min
  5. 5

    Why GANs Collapse

    Recognize mode collapse and name two common mitigations.

    12 min
  6. 6

    Destroying an Image on Purpose

    Add noise over a schedule and describe the image at each stage.

    12 min
  7. 7

    Learning to Undo One Step

    Explain what the network predicts at each denoising step.

    13 min
  8. 8

    Following the Data Uphill

    Read the score as a direction toward more probable images.

    12 min
  9. 9

    Sampling Schedules

    Trade steps against quality and pick a sampler for a budget.

    11 min
  10. 10

    Conditioning and Guidance

    Explain classifier-free guidance and predict what a high scale does.

    13 min
  11. 11

    Diffusion in Latent Space

    Explain why running diffusion on latents cut the cost so sharply.

    12 min
  12. 12

    From Words to Pictures

    Trace how a text encoder steers image generation, and where it misreads.

    12 min
  13. 13

    Making Time Consistent

    Explain what changes when frames must agree with each other.

    12 min
  14. 14

    Scoring a Picture Maker

    Read an FID number honestly and name three things it cannot see.

    12 min
  15. 15

    What These Models Understand

    Judge from failure cases what the model represents and what it fakes.

    12 min

Before you start

Keep going

All chapters