D7 · Neural Scaling Laws
The Kaplan Result
Summarize what the 2020 scaling paper claimed and how it was measured.
In January 2020, an OpenAI team made a bold offer: tell us your compute budget, and we will tell you your loss before you train. They trained transformers from 768 parameters up to 1.5 billion. Every run landed near the same straight line. Four months later, the same lab shipped GPT-3, over a hundred times bigger.