A chat model never picks a word. It rolls a weighted die.
F3 · Being Wrong on Purpose
Models do not output answers. They output distributions, and every choice after that is a sample. This chapter builds distributions from counting, then turns them into the two quantities you will use forever: cross-entropy, the loss almost every classifier and language model minimizes, and KL divergence, the extra bits a model pays for believing something other than the truth. Bayes' rule arrives as a belief update, not a formula to memorize. You leave knowing exactly what temperature does to a model's die.
- A Number You Do Not Know Yet
- Coins and Categories
- The Bell and the Flat Line
- Where It Sits and How Much It Wobbles
- Joint, Marginal, Conditional
- Bayes Is a Belief Update
- When Knowing One Tells You Nothing
- Sampling and the Law of Large Numbers
- Which Setting Explains the Data Best
- Surprise, Measured
- Cross-Entropy: The Loss You Will Use Forever
- KL Divergence
- Softmax and Temperature