R2 · Landmark Papers, Rebuilt
Chinchilla (2022)
Split one training budget the compute-optimal way, and say what it changed.
After 2020 the advice was to spend a bigger budget mostly on a bigger model. So the field trained hundreds of billions of weights on a few hundred billion tokens. The data barely moved from one model to the next.