F4 · Reading Data Honestly
Could Luck Have Done It
Test whether luck could explain the gap between two models, and say what the p-value does and does not mean.
Model A scored 82 on 200 questions. Model B scored 87. They answered 174 of them the same way. Everything you know about the gap lives in the other 26.