L9 · Models That Think Longer
Rewards You Can Check
Explain how a program that checks answers can train a model to reason, and how the group of attempts decides the push.
In L4, assistants were trained against a reward model, a network guessing what people prefer. It could be fooled. A math answer needs no guess. It matches the known answer or it does not.