Stage 2 · Intermediate · D5
Seeing
The moment we stopped understanding our own models.
13 lessons · 171 minSteady
About this chapter
A convolution is one small filter dragged across an image, and stacking them builds detectors for edges, then textures, then objects. This chapter draws every kernel so you can see what it fires on, then follows the path from LeNet to AlexNet to ResNet. You will see that AlexNet won in 2012 with four ordinary ingredients, not one big idea. It ends with adversarial examples, where a nudge too small to see turns a panda into a gibbon.
What you will be able to do
- 111 min
An Image Is a Tensor
Name the batch, channel, height and width axes of an image tensor and count what is inside.
- 213 min
One Filter, Dragged Across
Compute a convolution by hand and say what a bright spot in the output map means.
- 313 min
Kernels You Can Draw
Design a kernel that detects a pattern you choose, and predict what it will fire on.
- 413 min
Padding and Stride
Predict the output size from the input size, kernel, padding and stride, and choose them on purpose.
- 512 min
Pooling
Explain what max and average pooling keep, what they throw away, and when to use neither.
- 614 min
How Much Can This Neuron See?
Compute the receptive field of a unit deep in a stack, and say why downsampling is what makes it grow.
- 714 min
Channels and Depth
Count the parameters in a convolution layer and say what one output channel is made of.
- 814 min
What Actually Won in 2012
List AlexNet's ingredients and say which of them were new and which were borrowed.
- 914 min
Looking Inside a Vision Model
Read feature visualizations from early to late layers, and name what they cannot tell you.
- 1013 min
Going Deep With ResNet
Read a ResNet as stages of residual blocks, and price the bottleneck block that made 152 layers affordable.
- 1114 min
Reusing a Trained Model
Reuse a pretrained vision model on a small task, and choose what to freeze.
- 1213 min
When Attention Came for Vision
Compare a vision transformer with a convolutional network on built-in assumptions and hunger for data.
- 1313 min
A Panda, One Nudge, a Gibbon
Build an adversarial example, explain why it works, and say why defenses are hard.