Stage 3 · Expert · L9
Models That Think Longer
On many problems, letting a model think longer beats making it fourteen times bigger.
14 lessons · 135 minDemanding
About this chapter
A model that is right 95 times in 100 on every step still fails a ten-step problem four times in ten. This chapter follows the ways around that: writing the steps down, asking several times and taking the vote, checking each step, and training with rewards a program can verify. You will see how o1 and DeepSeek-R1 came out of those ideas, where thinking longer wastes money or tells a story that is not what happened, and how to read a reasoning score without being fooled.
What you will be able to do
- 18 min
One Guess Is Not Enough
Explain why a small chance of error on each step becomes a large one on a long problem.
- 29 min
Steps on the Page
Say what writing the steps changes, and why it first worked only in the largest models.
- 39 min
Ask Again, Take the Vote
Explain when a vote over many runs fixes an answer, and when it makes a wrong one more certain.
- 410 min
Thinking Time or a Bigger Model
Compare spending compute at answer time with training a bigger model, and say when each wins.
- 510 min
A Checker Picks the Best
Describe best-of-N with a checker, and explain why too many attempts can make it worse.
- 610 min
Grade the Answer or the Working
Tell an outcome reward from a process reward, and say what each one catches and costs.
- 711 min
Rewards You Can Check
Explain how a program that checks answers can train a model to reason, and how the group of attempts decides the push.
- 810 min
What the Training Grew
Describe what DeepSeek-R1-Zero learned from checkable rewards alone, and what that training did not add.
- 99 min
The Reasoning Models
Say what a reasoning model is, how it arrived, and what its thinking costs you.
- 1010 min
Searching a Tree of Thoughts
Run a beam search over partial solutions, and say what it buys and why it is hard to scale.
- 119 min
Tools in the Middle of a Thought
Explain why a reasoning model hands exact work to tools, and what a tool cannot fix.
- 129 min
Teaching Small Models to Think
Explain how a small model learns reasoning from a big model's worked chains, and why that can beat training it directly.
- 1310 min
When Thinking Goes Wrong
Recognize overthinking and unfaithful reasoning, and say what each costs you.
- 1411 min
Measuring Reasoning Honestly
Read a reasoning score critically: tries, test size, fresh questions and leaked ones.