Stage 1 · Beginner · D3
Backpropagation
The F = ma of artificial intelligence.
14 lessons · 152 minSteady
About this chapter
Gradient descent told you to step downhill. It never told you how a model with a hundred billion weights works out which way downhill is, and that is what this chapter does.
Backpropagation is the chain rule applied to a graph, and nothing more. You will draw the graph, push numbers forward, then push gradients back edge by edge until every weight knows its share of the blame. You will do one neuron by hand, then a two layer network by hand, then the same thing in matrices.
Then comes the honest part. Gradients fade through deep stacks, explode through recurrent ones, and die inside ReLU units. You will see each failure, measure it, and check every gradient you compute against a finite difference. You finish with the history, from a Finnish thesis in 1970 to Nature in 1986, and with the reasons neuroscientists think brains do something else.
What you will be able to do
- 111 min
Every Model Is a Graph
Draw any expression as a graph of small operations.
- 210 min
The Forward Pass
Push values through a graph and store what the backward pass will need.
- 312 min
Every Node Knows One Small Derivative
Write the local gradient for add, multiply and a nonlinearity.
- 412 min
Chaining Backwards
Multiply local gradients along a path to get any weight's gradient.
- 511 min
One Neuron, By Hand
Compute every gradient in a single sigmoid neuron with a pencil.
- 612 min
Two Layers, By Hand
Run a full backward pass through a two layer network by hand.
- 711 min
Backprop in Matrices
See why real libraries run the backward pass as a few matrix products, and check that every shape fits.
- 810 min
Reverse-Mode Autodiff
See why backprop walks the graph backwards: with one loss and millions of weights, one backward sweep gives every gradient.
- 99 min
Why It Costs So Little
Compare the cost of backprop to computing each gradient separately.
- 1010 min
Gradients That Fade
Explain why a gradient shrinks with depth, and measure how fast.
- 1110 min
Gradients That Explode, Units That Die
Spot exploding gradients and dead ReLUs from the numbers you log.
- 1212 min
Autograd, For Real
Train a small network with Masfofah autograd and read every gradient.
- 1311 min
Checking Gradients Numerically
Check your backward pass against a simple nudge test, and decide how close is close enough.
- 1411 min
What Backprop Is Not
State the main reasons the brain is unlikely to run backprop.