A model that scores 99% can be worse than a coin flip.
M5 · Did It Actually Work?
Most failed AI projects did not fail at training. They failed at evaluation, and nobody noticed until production. This chapter makes you paranoid in a useful way: start from a baseline, split data so the test set stays clean, and hunt leakage before you celebrate. You will run cross-validation, read learning curves, and do error analysis on the examples the model gets wrong. The habit you leave with is simple: know what would change your mind, and check it.
- The Score That Lies
- Beat the Dumbest Thing First
- One Metric Per Task
- Every Example Gets a Turn
- Splits That Respect Reality
- Too Stiff or Too Eager
- What the Curve Tells You to Do
- Searching Without Fooling Yourself
- The Wobble Between Runs
- Look at the Mistakes
- Hunting Data Leakage
- Benchmarks Wear Out
- Reporting Results Honestly