The fear that held neural networks back was aimed at the wrong thing.
D2 · Gradient Descent
Every model you have heard of learned the same way: feel which direction the ground tilts, take a step downhill, repeat a few million times. This chapter builds that loop from nothing. You will see loss as a landscape, watch a learning rate that is too large throw a model off a cliff, meet the fear of local minima that helped push many researchers away from neural networks, and see why high dimensions made that fear misplaced. Then the practical half: mini-batches and why noise helps, momentum, RMSProp, Adam and AdamW, warmup and cosine schedules, gradient clipping, and how to read a loss curve well enough to know what to change. You will build SGD, momentum and Adam by hand and race them…
- Loss Is a Landscape
- The Slope Tells You Which Way Is Down
- One Step Downhill
- Too Big, Too Small
- Bowls and Real Landscapes
- The Wrong Fear
- Noisy Steps Beat Perfect Ones
- Give the Ball Some Weight
- A Different Step for Every Weight
- Adam: Both Ideas at Once
- Change the Step as You Go
- The Spike That Ruins a Week
- Knowing When to Stop
- Watch Them Race