Skip to content

Stage 2 · Intermediate · L4

Training a Language Model

Pretraining never hands the model a single fact. It picks facts up because guessing the next word well is impossible without them.

14 lessons · 140 minDemanding

About this chapter

The objective is one sentence long: guess the next token, and pay for your surprise. Everything people call capability grows out of doing that on enough text. This chapter follows a real run from the first batch to the last checkpoint, then through the stages that turn a text predictor into an assistant. You will check a run against its bill, read a loss curve that goes wrong, and finish able to look at a benchmark table without being fooled by it.

What you will be able to do

  1. 1

    Guess the Next Token

    Say what a language model is trained to do, and why that one task makes it learn facts.

    9 min
  2. 2

    Grading the Guess

    Read a language model's loss as a perplexity, and say when two perplexities cannot be compared.

    10 min
  3. 3

    Where the Text Comes From

    Describe a real pretraining mix and say why the mix is not the pool.

    10 min
  4. 4

    Cleaning the Text

    Say what a web filter throws away and why duplicates are worse than junk.

    10 min
  5. 5

    Windows, Batches, Packing

    Explain how loose documents become a rectangle of tokens the model can read.

    9 min
  6. 6

    The Loop at Scale

    Walk one training step, and say what keeps time in a run that has no epochs.

    10 min
  7. 7

    Warmup, Then Decay

    Read a real run's learning-rate schedule, say why Adam needs a warmup, and say why some labs moved away from a cosine.

    10 min
  8. 8

    Small Numbers, Big Spikes

    Say why training uses two number formats at once, and what a team does when the loss spikes.

    10 min
  9. 9

    Splitting the Work

    Tell apart, by picture, the three ways a run is cut across many chips.

    10 min
  10. 10

    Tokens Against Parameters

    Check a published run against 6ND, and say which parameters count when a model is a mixture of experts.

    11 min
  11. 11

    From Predictor to Assistant

    Say what supervised fine-tuning changes, and what it leaves exactly as it was.

    10 min
  12. 12

    Learning From Preferences

    Describe how two replies and a human choice become a score, and why a leash is needed.

    11 min
  13. 13

    DPO, and Doing Less

    Compare DPO with RLHF on moving parts, and say where written rules replace human labels.

    10 min
  14. 14

    Reading a Benchmark Table

    Give three reasons to distrust a benchmark number, and test one of them yourself.

    10 min

Before you start

Keep going

All chapters