Skip to content

Stage 1 · Beginner · D2

Gradient Descent

The fear that held neural networks back was aimed at the wrong thing.

14 lessons · 169 minSteady

About this chapter

Every model you have heard of learned the same way: feel which direction the ground tilts, take a step downhill, repeat a few million times.

This chapter builds that loop from nothing. You will see loss as a landscape of hills and valleys, watch a learning rate that is too large throw a model off a cliff, meet the fear of local minima that helped push many researchers away from neural networks, and see why high dimensions made that fear misplaced.

Then the practical half: mini-batches and why noise helps, momentum, RMSProp, Adam and AdamW, warmup and cosine schedules, gradient clipping, and how to read a loss curve well enough to know what to change. You will build SGD, momentum and Adam by hand and race them on the same surface.

What you will be able to do

  1. 1

    Loss Is a Landscape

    Read a loss landscape: every setting of the weights is a place, and the loss is how high you stand.

    10 min
  2. 2

    The Slope Tells You Which Way Is Down

    Use the derivative at a point to decide which direction lowers the loss, and extend that to many weights at once.

    11 min
  3. 3

    One Step Downhill

    Write and run the gradient descent update, and explain what the learning rate controls.

    12 min
  4. 4

    Too Big, Too Small

    Recognize divergence and crawling from a loss curve, and know the range of learning rates that can work.

    12 min
  5. 5

    Bowls and Real Landscapes

    Tell a convex problem from a non-convex one, and say what each guarantees about where you end up.

    12 min
  6. 6

    The Wrong Fear

    Explain why local minima are not the real obstacle in high dimensions, and what a saddle point is instead.

    13 min
  7. 7

    Noisy Steps Beat Perfect Ones

    Choose between full-batch, single-example and mini-batch gradients, and explain why noise helps rather than hurts.

    13 min
  8. 8

    Give the Ball Some Weight

    Explain momentum as a running average of gradients, and predict where it helps and where it overshoots.

    12 min
  9. 9

    A Different Step for Every Weight

    Explain why one shared learning rate is a bad fit for real models, and how RMSProp gives each weight its own scale.

    12 min
  10. 10

    Adam: Both Ideas at Once

    See how Adam joins momentum and RMSProp, why its first steps need a fix, and what AdamW changes about weight decay.

    13 min
  11. 11

    Change the Step as You Go

    Pick a learning-rate schedule for a run, and say what warmup protects against.

    12 min
  12. 12

    The Spike That Ruins a Week

    Spot an exploding gradient, and cap its size without changing which way it points.

    12 min
  13. 13

    Knowing When to Stop

    Read a loss curve and decide whether to keep going, cut the learning rate, or stop.

    13 min
  14. 14

    Watch Them Race

    Race SGD, momentum, RMSProp and Adam on the same ground, and explain why the winner changes with the ground.

    12 min

Before you start

Keep going

All chapters