Skip to content

Stage 3 · Expert · L9

Models That Think Longer

On many problems, letting a model think longer beats making it fourteen times bigger.

14 lessons · 135 minDemanding

About this chapter

A model that is right 95 times in 100 on every step still fails a ten-step problem four times in ten. This chapter follows the ways around that: writing the steps down, asking several times and taking the vote, checking each step, and training with rewards a program can verify. You will see how o1 and DeepSeek-R1 came out of those ideas, where thinking longer wastes money or tells a story that is not what happened, and how to read a reasoning score without being fooled.

What you will be able to do

  1. 1

    One Guess Is Not Enough

    Explain why a small chance of error on each step becomes a large one on a long problem.

    8 min
  2. 2

    Steps on the Page

    Say what writing the steps changes, and why it first worked only in the largest models.

    9 min
  3. 3

    Ask Again, Take the Vote

    Explain when a vote over many runs fixes an answer, and when it makes a wrong one more certain.

    9 min
  4. 4

    Thinking Time or a Bigger Model

    Compare spending compute at answer time with training a bigger model, and say when each wins.

    10 min
  5. 5

    A Checker Picks the Best

    Describe best-of-N with a checker, and explain why too many attempts can make it worse.

    10 min
  6. 6

    Grade the Answer or the Working

    Tell an outcome reward from a process reward, and say what each one catches and costs.

    10 min
  7. 7

    Rewards You Can Check

    Explain how a program that checks answers can train a model to reason, and how the group of attempts decides the push.

    11 min
  8. 8

    What the Training Grew

    Describe what DeepSeek-R1-Zero learned from checkable rewards alone, and what that training did not add.

    10 min
  9. 9

    The Reasoning Models

    Say what a reasoning model is, how it arrived, and what its thinking costs you.

    9 min
  10. 10

    Searching a Tree of Thoughts

    Run a beam search over partial solutions, and say what it buys and why it is hard to scale.

    10 min
  11. 11

    Tools in the Middle of a Thought

    Explain why a reasoning model hands exact work to tools, and what a tool cannot fix.

    9 min
  12. 12

    Teaching Small Models to Think

    Explain how a small model learns reasoning from a big model's worked chains, and why that can beat training it directly.

    9 min
  13. 13

    When Thinking Goes Wrong

    Recognize overthinking and unfaithful reasoning, and say what each costs you.

    10 min
  14. 14

    Measuring Reasoning Honestly

    Read a reasoning score critically: tries, test size, fresh questions and leaked ones.

    11 min

Before you start

Keep going

All chapters