Stage 2 · Intermediate · M5
Did It Actually Work?
A model that scores 99% on its own examples can be worse than a coin flip on new ones.
13 lessons · 125 minSteady
About this chapter
A score is only as trustworthy as the test behind it. This chapter makes you careful in a useful way: start from a baseline, split data so the test set stays clean, and hunt leakage before you celebrate. You will run cross-validation, read learning curves, and do error analysis on the examples the model gets wrong. The habit you leave with is simple: know what would change your mind, and check it.
What you will be able to do
- 18 min
The Score That Lies
Say why a score measured on the training examples proves nothing, and what to hold back.
- 28 min
Beat the Dumbest Thing First
Build the trivial baseline for your task and record its score before anything else.
- 311 min
One Metric Per Task
Pick the one number your task is judged by, from the price of each kind of mistake.
- 410 min
Every Example Gets a Turn
Run k-fold cross-validation and report the mean with its spread instead of one number.
- 59 min
Splits That Respect Reality
Split by group or by time whenever a random split would let the model recognize what it already saw.
- 610 min
Too Stiff or Too Eager
Read two scores against your target and pick the fix that shrinks the bigger distance.
- 710 min
What the Curve Tells You to Do
Decide from a learning curve whether to add data, add flexibility, or stop.
- 811 min
Searching Without Fooling Yourself
Search over settings with three piles of data, and say how much the winner flatters you.
- 99 min
The Wobble Between Runs
Tell a real improvement from the wobble between seeds, and show that a part earns its place.
- 1010 min
Look at the Mistakes
Sort a sample of wrong answers into buckets and work out what each bucket is worth.
- 1111 min
Hunting Data Leakage
Run a leakage checklist over a pipeline and catch a column that will not exist at prediction time.
- 129 min
Benchmarks Wear Out
Explain how a shared test set loses its meaning, through reuse and through leaking into training data.
- 139 min
Reporting Results Honestly
Write a results row and a run record that someone who doubts you can check.