Stage 1 · Beginner · D4
Deep Learning
The geometry of depth.
14 lessons · 190 minSteady
About this chapter
A deep network arranges its neurons so each one works on what the last one already found.
This chapter is geometric. A layer stretches and bends space, a ReLU folds it, and folds compose: two layers of ten units cut the plane into far more pieces than one layer of twenty. You will see one hidden layer prove that width alone is enough, and then see the price of that proof, which is why nobody pays it.
Then come the parts that make depth trainable: where the weights start, why normalization steadies a stack, what dropout really costs, and why one skip connection let networks go from 20 layers to 152. You will train a multilayer network (an MLP) on two interleaved moons, read its curves, and decide when to stop. The last stop is double descent, a curve nobody fully explains yet.
What you will be able to do
- 114 min
Layers Fold Space
Watch a tangled dataset become separable one layer at a time.
- 213 min
Folds Made of ReLU
Explain a ReLU layer as a fold and count the pieces it creates.
- 313 min
One Layer Can Do Anything
Say what universal approximation promises, and what it leaves out.
- 414 min
Depth Is Cheaper Than Width
Show that stacking folds multiplies the pieces while spreading them only adds.
- 513 min
Features Nobody Designed
Read a hidden layer as a set of coordinates the network invented for itself.
- 612 min
Where the Weights Start Matters
Say why the starting size of the weights matters, and how Xavier and He keep the signal steady.
- 714 min
Batch Norm and Layer Norm
Apply the right normalization for a given architecture and batch size.
- 813 min
Dropout and Weight Decay
Pick a regularizer from the symptom you see in the curves.
- 914 min
The Shortcut That Opened Up Depth
Explain why adding the input back makes a very deep stack trainable.
- 1014 min
Five Lines, In Order
Write the training loop from memory and say what each line owns.
- 1114 min
Memorizing Versus Understanding
Detect memorization and choose between more data, more noise or less model.
- 1213 min
Double Descent
See why a model can get worse as it grows, and then better again.
- 1315 min
Build an MLP on Moons
Train a network with two hidden layers on moons on this device and tune it.
- 1414 min
Reading the Curves
Diagnose a run from its two curves and stop it at the right epoch.