Skip to content

Stage 2 · Intermediate · M5

Did It Actually Work?

A model that scores 99% on its own examples can be worse than a coin flip on new ones.

13 lessons · 125 minSteady

About this chapter

A score is only as trustworthy as the test behind it. This chapter makes you careful in a useful way: start from a baseline, split data so the test set stays clean, and hunt leakage before you celebrate. You will run cross-validation, read learning curves, and do error analysis on the examples the model gets wrong. The habit you leave with is simple: know what would change your mind, and check it.

What you will be able to do

  1. 1

    The Score That Lies

    Say why a score measured on the training examples proves nothing, and what to hold back.

    8 min
  2. 2

    Beat the Dumbest Thing First

    Build the trivial baseline for your task and record its score before anything else.

    8 min
  3. 3

    One Metric Per Task

    Pick the one number your task is judged by, from the price of each kind of mistake.

    11 min
  4. 4

    Every Example Gets a Turn

    Run k-fold cross-validation and report the mean with its spread instead of one number.

    10 min
  5. 5

    Splits That Respect Reality

    Split by group or by time whenever a random split would let the model recognize what it already saw.

    9 min
  6. 6

    Too Stiff or Too Eager

    Read two scores against your target and pick the fix that shrinks the bigger distance.

    10 min
  7. 7

    What the Curve Tells You to Do

    Decide from a learning curve whether to add data, add flexibility, or stop.

    10 min
  8. 8

    Searching Without Fooling Yourself

    Search over settings with three piles of data, and say how much the winner flatters you.

    11 min
  9. 9

    The Wobble Between Runs

    Tell a real improvement from the wobble between seeds, and show that a part earns its place.

    9 min
  10. 10

    Look at the Mistakes

    Sort a sample of wrong answers into buckets and work out what each bucket is worth.

    10 min
  11. 11

    Hunting Data Leakage

    Run a leakage checklist over a pipeline and catch a column that will not exist at prediction time.

    11 min
  12. 12

    Benchmarks Wear Out

    Explain how a shared test set loses its meaning, through reuse and through leaking into training data.

    9 min
  13. 13

    Reporting Results Honestly

    Write a results row and a run record that someone who doubts you can check.

    9 min

Before you start

Keep going

All chapters