Fourteen papers built the field. You can rebuild the core of each one.
R2 · Landmark Papers, Rebuilt
Reading a landmark paper is good. Rebuilding its central idea is better, because it forces every hidden assumption into the open. Each lesson here takes one paper, shows the problem that made it necessary, puts its key figure back together as something you can push on, and then says plainly what has not aged well. Where a full reproduction is impossible inside a lesson, you build the smallest honest version and we name what was lost. You end with a chain you can walk from 1958 to the models shipping today.
- Perceptron (1958)
- Learning Representations by Back-Propagating Errors (1986)
- LeNet (1998)
- AlexNet (2012)
- Word2Vec (2013)
- Sequence to Sequence (2014)
- Attention Is All You Need (2017)
- GPT-2 and GPT-3
- Scaling Laws (2020)
- Chinchilla (2022)
- CLIP (2021)
- Denoising Diffusion (2020)
- InstructGPT (2022)
- DeepSeek V2 and V3