The geometry of depth.
D4 · Deep Learning
A deep network arranges its neurons so each one works on what the last one already found. This chapter is geometric. A layer stretches and bends space, a ReLU folds it, and folds compose: two layers of ten units cut the plane into far more pieces than one layer of twenty. You will see one hidden layer prove that width alone is enough, and then see the price of that proof, which is why nobody pays it. Then the machinery that makes depth trainable: where the weights start, why normalization steadies a stack, what dropout really costs, and why one skip connection let networks go from 20 layers to 152. You will train an MLP on two interleaved moons, read its curves, and decide when to stop. The…
- Layers Fold Space
- Folds Made of ReLU
- One Layer Can Do Anything
- Depth Is Cheaper Than Width
- Features Nobody Designed
- Where the Weights Start Matters
- Batch Norm and Layer Norm
- Dropout and Weight Decay
- The Shortcut That Opened Up Depth
- Five Lines, In Order
- Memorizing Versus Understanding
- Double Descent
- Build an MLP on Moons
- Reading the Curves