L5 · Talking to Models
Room to Think
Say when working out loud buys accuracy and when it only buys tokens.
The model runs the same fixed work for every token it writes. No more for a hard question, no less for an easy one. So an answer given in one token gets one token's worth of work.