Stage 3 · Expert · L10
Smaller, Faster, Cheaper
Most of the time a chat model spends writing to you, its chip is waiting for memory.
14 lessons · 146 minResearch
About this chapter
A model nobody can afford to run helps nobody. This chapter is about the price of one answer, in memory, in time and in money, and every honest way to lower it. You will round weights to four bits and see what breaks, let a big model teach a small one, cut weights away, serve a thousand customers from one base model, route each token through a few experts, let a small model draft while a big one checks, and fit a model on a phone. Then you will choose between them the only way that holds up, with a measured eval, in Arabic as well as English.
What you will be able to do
- 19 min
Who Gets the Answer
Work out whether a model fits a device from its weights and its bits, and name the four ways to shrink it.
- 210 min
The Wall Is Memory
Explain why writing one token is limited by memory speed, and put a ceiling on tokens a second.
- 311 min
One Scale or Many
Round weights to 4 bits with one scale per group, and count what the scales cost.
- 411 min
One Number Ruins the Row
Show how one outlier wrecks rounding, and name two ways real methods protect the rest.
- 512 min
Round, Then Repair
Round weights one at a time and cancel each error with the weights still to come.
- 611 min
A Teacher With a Vocabulary
Train a student on a teacher's whole row of chances, and use heat to make the small ones teach.
- 711 min
Cutting Weights
Prune a layer three ways and say which pattern a chip can actually skip.
- 810 min
One Base, Many Adapters
Size an adapter, and serve many customers from one shared model.
- 910 min
Cheap to Run, Costly to Hold
Say what a mixture of experts saves and what it does not, on a busy server and on a phone.
- 1010 min
A Small Model Drafts
Explain how a big model checks a small model's draft in one pass, and why the text does not change.
- 1110 min
How Far Ahead to Guess
Work out the tokens one big step yields from a draft, and pick how many tokens to draft.
- 1210 min
The Long Conversation Bill
Stack the levers that shrink a long conversation's cache, and say what each one gives up.
- 1310 min
A Model in Your Pocket
Fit a language model into a phone's share of memory, and say what the tokenizer does to an Arabic reply.
- 1411 min
Choosing With a Measured Eval
Pick the cheapest setting that meets a bar set in advance, measured in every language your users speak.