Stage 1 · Beginner · D2
Gradient Descent
The fear that held neural networks back was aimed at the wrong thing.
14 lessons · 169 minSteady
About this chapter
Every model you have heard of learned the same way: feel which direction the ground tilts, take a step downhill, repeat a few million times.
This chapter builds that loop from nothing. You will see loss as a landscape of hills and valleys, watch a learning rate that is too large throw a model off a cliff, meet the fear of local minima that helped push many researchers away from neural networks, and see why high dimensions made that fear misplaced.
Then the practical half: mini-batches and why noise helps, momentum, RMSProp, Adam and AdamW, warmup and cosine schedules, gradient clipping, and how to read a loss curve well enough to know what to change. You will build SGD, momentum and Adam by hand and race them on the same surface.
What you will be able to do
- 110 min
Loss Is a Landscape
Read a loss landscape: every setting of the weights is a place, and the loss is how high you stand.
- 211 min
The Slope Tells You Which Way Is Down
Use the derivative at a point to decide which direction lowers the loss, and extend that to many weights at once.
- 312 min
One Step Downhill
Write and run the gradient descent update, and explain what the learning rate controls.
- 412 min
Too Big, Too Small
Recognize divergence and crawling from a loss curve, and know the range of learning rates that can work.
- 512 min
Bowls and Real Landscapes
Tell a convex problem from a non-convex one, and say what each guarantees about where you end up.
- 613 min
The Wrong Fear
Explain why local minima are not the real obstacle in high dimensions, and what a saddle point is instead.
- 713 min
Noisy Steps Beat Perfect Ones
Choose between full-batch, single-example and mini-batch gradients, and explain why noise helps rather than hurts.
- 812 min
Give the Ball Some Weight
Explain momentum as a running average of gradients, and predict where it helps and where it overshoots.
- 912 min
A Different Step for Every Weight
Explain why one shared learning rate is a bad fit for real models, and how RMSProp gives each weight its own scale.
- 1013 min
Adam: Both Ideas at Once
See how Adam joins momentum and RMSProp, why its first steps need a fix, and what AdamW changes about weight decay.
- 1112 min
Change the Step as You Go
Pick a learning-rate schedule for a run, and say what warmup protects against.
- 1212 min
The Spike That Ruins a Week
Spot an exploding gradient, and cap its size without changing which way it points.
- 1313 min
Knowing When to Stop
Read a loss curve and decide whether to keep going, cut the learning rate, or stop.
- 1412 min
Watch Them Race
Race SGD, momentum, RMSProp and Adam on the same ground, and explain why the winner changes with the ground.