The moment we stopped understanding our own models.
D5 · Seeing
A convolution is one small filter dragged across an image, and stacking them builds detectors for edges, then textures, then objects. This chapter draws every kernel so you can see what it fires on, then follows the path from LeNet to AlexNet to ResNet. You will learn why AlexNet won in 2012, and it was four ordinary ingredients rather than one idea. It ends with adversarial examples, where a nudge too small to see turns a panda into a gibbon.
- An Image Is a Tensor
- One Filter, Dragged Across
- Kernels You Can Draw
- Padding and Stride
- Pooling
- How Much Can This Neuron See?
- Channels and Depth
- What Actually Won in 2012
- Looking Inside a Vision Model
- Going Deep With ResNet
- Reusing a Trained Model
- When Attention Came for Vision
- A Panda, One Nudge, a Gibbon