AI cannot cross this line, and nobody knows why the line is there.
D7 · Neural Scaling Laws
Make a model bigger, feed it more data, spend more compute, and the loss falls along a straight line on a log-log plot. That regularity is one of the strangest empirical facts in the field, and it now decides how billions of dollars are spent. This chapter teaches you to read those plots, compare the Kaplan and Chinchilla recipes, and budget compute between parameters and tokens. It is also honest about emergence, where the evidence is contested.
- Reading a Log-Log Plot
- Power Laws Everywhere
- Compute, Data, Parameters
- The Kaplan Result
- Chinchilla Changes the Recipe
- Spending a Compute Budget
- Emergence and Its Critics
- Why Bigger Keeps Working
- What a Frontier Run Costs
- Scaling at Inference Time
- What the Curve Does Not Predict