Most of the time a chat model spends writing to you, its chip is waiting for memory.
L10 · Smaller, Faster, Cheaper
A model nobody can afford to run helps nobody. This chapter is about the price of one answer, in memory, in time and in money, and every honest way to lower it. You will round weights to four bits and see what breaks, let a big model teach a small one, cut weights away, serve a thousand customers from one base model, route each token through a few experts, let a small model draft while a big one checks, and fit a model on a phone. Then you will choose between them the only way that holds up: with a measured eval, in Arabic as well as English.
- Who Gets the Answer
- The Wall Is Memory
- One Scale or Many
- One Number Ruins the Row
- Round, Then Repair
- A Teacher With a Vocabulary
- Cutting Weights
- One Base, Many Adapters
- Cheap to Run, Costly to Hold
- A Small Model Drafts
- How Far Ahead to Guess
- The Long Conversation Bill
- A Model in Your Pocket
- Choosing With a Measured Eval