Skip to content

Stage 2 · Intermediate · F4

Reading Data Honestly

Train the same model twice with different random seeds and you get two scores. Some published wins are smaller than that gap.

12 lessons · 114 minSteady

About this chapter

Statistics is how you tell a real improvement from a lucky one. Each lesson asks a plain question about a number: where is the middle, how far does it wobble, who is missing from the sample, and could luck alone have done this. You will draw confidence intervals, build one by resampling your own test set, run a coin test on two models, and see how twenty tries and early peeks manufacture wins. The chapter ends on a results table taken apart run by run.

What you will be able to do

  1. 1

    Nobody Is Average

    Choose between mean and median, and report spread with the right number.

    7 min
  2. 2

    Shapes Real Data Takes

    Recognize skew, long tails and multiple peaks in real data, and say what each breaks.

    9 min
  3. 3

    Two Columns That Move Together

    Read a correlation coefficient and name a plausible confounder for a given pair.

    9 min
  4. 4

    Who Is Missing From the Sample

    Spot selection and survivorship bias in a dataset description.

    9 min
  5. 5

    Luck Has a Size

    Say how far a score can drift on luck alone, and how much data it takes to halve that drift.

    9 min
  6. 6

    The Interval Is the Result

    Report a score as an interval, and say what the ninety five percent actually counts.

    9 min
  7. 7

    Pull the Test Set Up by Itself

    Resample one test set to get an interval for any score you can compute.

    10 min
  8. 8

    Could Luck Have Done It

    Test whether luck could explain the gap between two models, and say what the p-value does and does not mean.

    11 min
  9. 9

    Right Nineteen Times in Twenty

    Work out how often a flag, or a result under 0.05, is right when the real thing is rare.

    10 min
  10. 10

    Keep Trying and You Will Win

    Explain how repeated tries manufacture a result, and correct for the number of tries.

    10 min
  11. 11

    Decide the Size Before You Look

    Design an A/B test with a sample size fixed in advance, and say why peeking breaks it.

    10 min
  12. 12

    Reading a Table With Suspicion

    Decide whether a reported gain is worth believing, and name what would make you believe it.

    11 min

Before you start

Keep going

All chapters