Stage 2 · Intermediate · L4
Training a Language Model
Pretraining never hands the model a single fact. It picks facts up because guessing the next word well is impossible without them.
14 lessons · 140 minDemanding
About this chapter
The objective is one sentence long: guess the next token, and pay for your surprise. Everything people call capability grows out of doing that on enough text. This chapter follows a real run from the first batch to the last checkpoint, then through the stages that turn a text predictor into an assistant. You will check a run against its bill, read a loss curve that goes wrong, and finish able to look at a benchmark table without being fooled by it.
What you will be able to do
- 19 min
Guess the Next Token
Say what a language model is trained to do, and why that one task makes it learn facts.
- 210 min
Grading the Guess
Read a language model's loss as a perplexity, and say when two perplexities cannot be compared.
- 310 min
Where the Text Comes From
Describe a real pretraining mix and say why the mix is not the pool.
- 410 min
Cleaning the Text
Say what a web filter throws away and why duplicates are worse than junk.
- 59 min
Windows, Batches, Packing
Explain how loose documents become a rectangle of tokens the model can read.
- 610 min
The Loop at Scale
Walk one training step, and say what keeps time in a run that has no epochs.
- 710 min
Warmup, Then Decay
Read a real run's learning-rate schedule, say why Adam needs a warmup, and say why some labs moved away from a cosine.
- 810 min
Small Numbers, Big Spikes
Say why training uses two number formats at once, and what a team does when the loss spikes.
- 910 min
Splitting the Work
Tell apart, by picture, the three ways a run is cut across many chips.
- 1011 min
Tokens Against Parameters
Check a published run against 6ND, and say which parameters count when a model is a mixture of experts.
- 1110 min
From Predictor to Assistant
Say what supervised fine-tuning changes, and what it leaves exactly as it was.
- 1211 min
Learning From Preferences
Describe how two replies and a human choice become a score, and why a leash is needed.
- 1310 min
DPO, and Doing Less
Compare DPO with RLHF on moving parts, and say where written rules replace human labels.
- 1410 min
Reading a Benchmark Table
Give three reasons to distrust a benchmark number, and test one of them yourself.