Skip to content

Stage 2 · Intermediate · D5

Seeing

The moment we stopped understanding our own models.

13 lessons · 171 minSteady

About this chapter

A convolution is one small filter dragged across an image, and stacking them builds detectors for edges, then textures, then objects. This chapter draws every kernel so you can see what it fires on, then follows the path from LeNet to AlexNet to ResNet. You will see that AlexNet won in 2012 with four ordinary ingredients, not one big idea. It ends with adversarial examples, where a nudge too small to see turns a panda into a gibbon.

What you will be able to do

  1. 1

    An Image Is a Tensor

    Name the batch, channel, height and width axes of an image tensor and count what is inside.

    11 min
  2. 2

    One Filter, Dragged Across

    Compute a convolution by hand and say what a bright spot in the output map means.

    13 min
  3. 3

    Kernels You Can Draw

    Design a kernel that detects a pattern you choose, and predict what it will fire on.

    13 min
  4. 4

    Padding and Stride

    Predict the output size from the input size, kernel, padding and stride, and choose them on purpose.

    13 min
  5. 5

    Pooling

    Explain what max and average pooling keep, what they throw away, and when to use neither.

    12 min
  6. 6

    How Much Can This Neuron See?

    Compute the receptive field of a unit deep in a stack, and say why downsampling is what makes it grow.

    14 min
  7. 7

    Channels and Depth

    Count the parameters in a convolution layer and say what one output channel is made of.

    14 min
  8. 8

    What Actually Won in 2012

    List AlexNet's ingredients and say which of them were new and which were borrowed.

    14 min
  9. 9

    Looking Inside a Vision Model

    Read feature visualizations from early to late layers, and name what they cannot tell you.

    14 min
  10. 10

    Going Deep With ResNet

    Read a ResNet as stages of residual blocks, and price the bottleneck block that made 152 layers affordable.

    13 min
  11. 11

    Reusing a Trained Model

    Reuse a pretrained vision model on a small task, and choose what to freeze.

    14 min
  12. 12

    When Attention Came for Vision

    Compare a vision transformer with a convolutional network on built-in assumptions and hunger for data.

    13 min
  13. 13

    A Panda, One Nudge, a Gibbon

    Build an adversarial example, explain why it works, and say why defenses are hard.

    13 min

Before you start

Keep going

All chapters