Stage 3 · Expert · R2
Landmark Papers, Rebuilt
Fourteen papers built the field. You can rebuild the core of each one.
14 lessons · 154 minResearch
About this chapter
Reading a landmark paper is good. Rebuilding its central idea is better, because it forces every hidden assumption into the open. Each lesson here takes one paper, shows the problem that made it necessary, puts its key figure back together as something you can push on, and then says plainly what has not aged well. Where a full reproduction is impossible inside a lesson, you build the smallest honest version and we name what was lost. You end with a chain you can walk from 1958 to the models shipping today.
What you will be able to do
- 111 min
Perceptron (1958)
Run the 1958 learning rule, and say exactly what it can and cannot separate.
- 212 min
Learning Representations by Back-Propagating Errors (1986)
Send an error backward through a middle layer and watch XOR fall.
- 311 min
LeNet (1998)
Count what weight sharing saves, and say what LeNet-5 reported on digits.
- 411 min
AlexNet (2012)
Say what AlexNet actually changed, and where its sixty million weights sat.
- 510 min
Word2Vec (2013)
Build word vectors from company alone, and test the analogy claim honestly.
- 610 min
Sequence to Sequence (2014)
Explain the fixed-vector bottleneck, and why reading the source backward helped.
- 713 min
Attention Is All You Need (2017)
Compute one attention head by hand, and say why the block replaced recurrence.
- 811 min
GPT-2 and GPT-3
Say what changed between the two, and what in-context learning does and does not do.
- 910 min
Scaling Laws (2020)
Fit a power law to a sweep of your own, and read what the exponent promises.
- 1010 min
Chinchilla (2022)
Split one training budget the compute-optimal way, and say what it changed.
- 1111 min
CLIP (2021)
Build the image and text matching matrix, and say what zero-shot really means.
- 1212 min
Denoising Diffusion (2020)
Follow a picture into noise and back, and say what the network is trained to predict.
- 1311 min
InstructGPT (2022)
Fit a reward model from comparisons, and say what human preference data bought.
- 1411 min
DeepSeek V2 and V3
Price the attention cache and a sparse model, and say what each choice bought.