Stage 1 · Beginner · F3
Being Wrong on Purpose
A chat model doesn't pick its next word outright. It rolls a weighted die.
13 lessons · 126 minSteady
About this chapter
A model's output is a distribution, not an answer, and every choice after that is a sample. This chapter builds distributions from counting, then turns them into the two quantities you will keep using: cross-entropy, the loss almost every classifier and language model minimizes, and KL divergence, the extra bits a model pays for believing something other than the truth. Bayes' rule arrives as a belief update, not a formula to memorize. You leave knowing exactly what temperature does to a model's die.
What you will be able to do
- 17 min
A Number You Do Not Know Yet
Describe an uncertain quantity as a random variable with a set of outcomes.
- 29 min
Coins and Categories
Use Bernoulli and categorical distributions and read their parameters.
- 310 min
The Bell and the Flat Line
Read a bell curve and a flat one, and say why the area under a curve, not its height, is the chance.
- 49 min
Where It Sits and How Much It Wobbles
Find what a game pays on average, and say why two games with the same average can feel very different.
- 510 min
Joint, Marginal, Conditional
Move between joint, marginal and conditional probabilities on one table.
- 612 min
Bayes Is a Belief Update
Explain why a positive test for a rare disease is usually a false alarm, and how new evidence moves a belief.
- 77 min
When Knowing One Tells You Nothing
Test independence and spot the wrong independence assumptions in real models.
- 810 min
Sampling and the Law of Large Numbers
Estimate a chance by sampling, and say how much more data it takes to cut the error in half.
- 910 min
Which Setting Explains the Data Best
Find the coin setting that explains some flips best, and say how training a model does the same thing.
- 1010 min
Surprise, Measured
Measure how surprising a forecast is, in bits, and say which of two is harder to predict.
- 1111 min
Cross-Entropy: The Loss You Will Keep Using
Compute cross-entropy for one prediction and explain why a confident wrong answer costs so much.
- 1210 min
KL Divergence
Say what a wrong belief costs in extra bits, and why swapping the two beliefs changes the number.
- 1311 min
Softmax and Temperature
Turn scores into probabilities and predict the effect of raising or lowering temperature.