L4 · Training a Language Model
Reading a Benchmark Table
Give three reasons to distrust a benchmark number, and test one of them yourself.
The best-known test of what a model knows has 15,908 multiple-choice questions in 57 subjects, from law to astronomy. When it came out, GPT-3 averaged 44 out of 100. That one number hid a spread from 69 down to 26.