L4 · Training a Language Model
Tokens Against Parameters
Check a published run against 6ND, and say which parameters count when a model is a mixture of experts.
The scaling laws chapter found the cheapest split of a budget: about 20 tokens for every parameter. This lesson puts real runs next to that rule, and next to , to see which claims hold up.