The F = ma of artificial intelligence.
D3 · Backpropagation
Gradient descent told you to step downhill. It never told you how a model with a hundred billion weights works out which way downhill is. That is this chapter. Backpropagation is the chain rule applied to a graph, and nothing more. You will draw the graph, push numbers forward, then push gradients back edge by edge until every weight knows its share of the blame. You will do one neuron by hand, then a two layer network by hand, then the same thing in matrices. Then the honest part. Gradients fade through deep stacks, explode through recurrent ones, and die inside ReLU units. You will see each failure, measure it, and check every gradient you compute against a finite difference. You finish…
- Every Model Is a Graph
- The Forward Pass
- Every Node Knows One Small Derivative
- Chaining Backwards
- One Neuron, By Hand
- Two Layers, By Hand
- Backprop in Matrices
- Reverse-Mode Autodiff
- Why It Costs So Little
- Gradients That Fade
- Gradients That Explode, Units That Die
- Autograd, For Real
- Checking Gradients Numerically
- What Backprop Is Not