Skip to content

Stage 1 · Beginner · F3

Being Wrong on Purpose

A chat model doesn't pick its next word outright. It rolls a weighted die.

13 lessons · 126 minSteady

About this chapter

A model's output is a distribution, not an answer, and every choice after that is a sample. This chapter builds distributions from counting, then turns them into the two quantities you will keep using: cross-entropy, the loss almost every classifier and language model minimizes, and KL divergence, the extra bits a model pays for believing something other than the truth. Bayes' rule arrives as a belief update, not a formula to memorize. You leave knowing exactly what temperature does to a model's die.

What you will be able to do

  1. 1

    A Number You Do Not Know Yet

    Describe an uncertain quantity as a random variable with a set of outcomes.

    7 min
  2. 2

    Coins and Categories

    Use Bernoulli and categorical distributions and read their parameters.

    9 min
  3. 3

    The Bell and the Flat Line

    Read a bell curve and a flat one, and say why the area under a curve, not its height, is the chance.

    10 min
  4. 4

    Where It Sits and How Much It Wobbles

    Find what a game pays on average, and say why two games with the same average can feel very different.

    9 min
  5. 5

    Joint, Marginal, Conditional

    Move between joint, marginal and conditional probabilities on one table.

    10 min
  6. 6

    Bayes Is a Belief Update

    Explain why a positive test for a rare disease is usually a false alarm, and how new evidence moves a belief.

    12 min
  7. 7

    When Knowing One Tells You Nothing

    Test independence and spot the wrong independence assumptions in real models.

    7 min
  8. 8

    Sampling and the Law of Large Numbers

    Estimate a chance by sampling, and say how much more data it takes to cut the error in half.

    10 min
  9. 9

    Which Setting Explains the Data Best

    Find the coin setting that explains some flips best, and say how training a model does the same thing.

    10 min
  10. 10

    Surprise, Measured

    Measure how surprising a forecast is, in bits, and say which of two is harder to predict.

    10 min
  11. 11

    Cross-Entropy: The Loss You Will Keep Using

    Compute cross-entropy for one prediction and explain why a confident wrong answer costs so much.

    11 min
  12. 12

    KL Divergence

    Say what a wrong belief costs in extra bits, and why swapping the two beliefs changes the number.

    10 min
  13. 13

    Softmax and Temperature

    Turn scores into probabilities and predict the effect of raising or lowering temperature.

    11 min

Before you start

Keep going

All chapters