Skip to content

Stage 1 · Beginner · D3

Backpropagation

The F = ma of artificial intelligence.

14 lessons · 152 minSteady

About this chapter

Gradient descent told you to step downhill. It never told you how a model with a hundred billion weights works out which way downhill is, and that is what this chapter does.

Backpropagation is the chain rule applied to a graph, and nothing more. You will draw the graph, push numbers forward, then push gradients back edge by edge until every weight knows its share of the blame. You will do one neuron by hand, then a two layer network by hand, then the same thing in matrices.

Then comes the honest part. Gradients fade through deep stacks, explode through recurrent ones, and die inside ReLU units. You will see each failure, measure it, and check every gradient you compute against a finite difference. You finish with the history, from a Finnish thesis in 1970 to Nature in 1986, and with the reasons neuroscientists think brains do something else.

What you will be able to do

  1. 1

    Every Model Is a Graph

    Draw any expression as a graph of small operations.

    11 min
  2. 2

    The Forward Pass

    Push values through a graph and store what the backward pass will need.

    10 min
  3. 3

    Every Node Knows One Small Derivative

    Write the local gradient for add, multiply and a nonlinearity.

    12 min
  4. 4

    Chaining Backwards

    Multiply local gradients along a path to get any weight's gradient.

    12 min
  5. 5

    One Neuron, By Hand

    Compute every gradient in a single sigmoid neuron with a pencil.

    11 min
  6. 6

    Two Layers, By Hand

    Run a full backward pass through a two layer network by hand.

    12 min
  7. 7

    Backprop in Matrices

    See why real libraries run the backward pass as a few matrix products, and check that every shape fits.

    11 min
  8. 8

    Reverse-Mode Autodiff

    See why backprop walks the graph backwards: with one loss and millions of weights, one backward sweep gives every gradient.

    10 min
  9. 9

    Why It Costs So Little

    Compare the cost of backprop to computing each gradient separately.

    9 min
  10. 10

    Gradients That Fade

    Explain why a gradient shrinks with depth, and measure how fast.

    10 min
  11. 11

    Gradients That Explode, Units That Die

    Spot exploding gradients and dead ReLUs from the numbers you log.

    10 min
  12. 12

    Autograd, For Real

    Train a small network with Masfofah autograd and read every gradient.

    12 min
  13. 13

    Checking Gradients Numerically

    Check your backward pass against a simple nudge test, and decide how close is close enough.

    11 min
  14. 14

    What Backprop Is Not

    State the main reasons the brain is unlikely to run backprop.

    11 min

Before you start

Keep going

All chapters