Skip to content

Stage 3 · Expert · R2

Landmark Papers, Rebuilt

Fourteen papers built the field. You can rebuild the core of each one.

14 lessons · 154 minResearch

About this chapter

Reading a landmark paper is good. Rebuilding its central idea is better, because it forces every hidden assumption into the open. Each lesson here takes one paper, shows the problem that made it necessary, puts its key figure back together as something you can push on, and then says plainly what has not aged well. Where a full reproduction is impossible inside a lesson, you build the smallest honest version and we name what was lost. You end with a chain you can walk from 1958 to the models shipping today.

What you will be able to do

  1. 1

    Perceptron (1958)

    Run the 1958 learning rule, and say exactly what it can and cannot separate.

    11 min
  2. 2

    Learning Representations by Back-Propagating Errors (1986)

    Send an error backward through a middle layer and watch XOR fall.

    12 min
  3. 3

    LeNet (1998)

    Count what weight sharing saves, and say what LeNet-5 reported on digits.

    11 min
  4. 4

    AlexNet (2012)

    Say what AlexNet actually changed, and where its sixty million weights sat.

    11 min
  5. 5

    Word2Vec (2013)

    Build word vectors from company alone, and test the analogy claim honestly.

    10 min
  6. 6

    Sequence to Sequence (2014)

    Explain the fixed-vector bottleneck, and why reading the source backward helped.

    10 min
  7. 7

    Attention Is All You Need (2017)

    Compute one attention head by hand, and say why the block replaced recurrence.

    13 min
  8. 8

    GPT-2 and GPT-3

    Say what changed between the two, and what in-context learning does and does not do.

    11 min
  9. 9

    Scaling Laws (2020)

    Fit a power law to a sweep of your own, and read what the exponent promises.

    10 min
  10. 10

    Chinchilla (2022)

    Split one training budget the compute-optimal way, and say what it changed.

    10 min
  11. 11

    CLIP (2021)

    Build the image and text matching matrix, and say what zero-shot really means.

    11 min
  12. 12

    Denoising Diffusion (2020)

    Follow a picture into noise and back, and say what the network is trained to predict.

    12 min
  13. 13

    InstructGPT (2022)

    Fit a reward model from comparisons, and say what human preference data bought.

    11 min
  14. 14

    DeepSeek V2 and V3

    Price the attention cache and a sparse model, and say what each choice bought.

    11 min

Before you start

Keep going

All chapters