Stage 2 · Intermediate · D7
Neural Scaling Laws
AI cannot cross this line, and nobody knows why the line is there.
11 lessons · 110 minDemanding
About this chapter
Make a model bigger, feed it more data, spend more compute, and the loss falls along a straight line on a log-log plot. That regularity is one of the strangest empirical facts in the field, and it now decides how billions of dollars are spent. This chapter teaches you to read those plots, compare the Kaplan and Chinchilla recipes, and budget compute between parameters and tokens. It is also honest about emergence, where the evidence is contested.
What you will be able to do
- 110 min
Reading a Log-Log Plot
Read a slope on log axes and turn it back into a plain-language claim.
- 210 min
Power Laws Everywhere
Recognize a power law and state what it predicts and for how long.
- 39 min
Compute, Data, Parameters
Explain how the three inputs trade against each other for a fixed loss.
- 410 min
The Kaplan Result
Summarize what the 2020 scaling paper claimed and how it was measured.
- 511 min
Chinchilla Changes the Recipe
Explain why models were undertrained and what the corrected ratio is.
- 610 min
Spending a Compute Budget
Split a fixed budget between model size and tokens, with a reason.
- 711 min
Emergence and Its Critics
Explain how metric choice can manufacture a sudden jump in ability.
- 810 min
Why Bigger Keeps Working
Give the leading explanations for scaling and mark which are still guesses.
- 910 min
What a Frontier Run Costs
Estimate the compute and money behind a large training run.
- 109 min
Scaling at Inference Time
Compare spending compute during training against spending it per answer.
- 1110 min
What the Curve Does Not Predict
List the things loss curves say nothing about, including data running out.