D7 · Neural Scaling Laws
Chinchilla Changes the Recipe
Explain why models were undertrained and what the corrected ratio is.
In March 2022, DeepMind trained Chinchilla: 70 billion parameters. Their earlier model, Gopher, had 280 billion. Both runs cost the same. Chinchilla won anyway, 67.5 percent on the MMLU knowledge test against Gopher's 60.0. The trick: it read almost five times as much text.