L9 · Models That Think Longer
What the Training Grew
Describe what DeepSeek-R1-Zero learned from checkable rewards alone, and what that training did not add.
R1-Zero started from a base model and was never fine-tuned on worked solutions. It only had lesson 7's two rules. On a hard math contest, its score on a single try rose from 15.6 to 71.0 percent.