Skip to content

Stage 3 · Expert · L10

Smaller, Faster, Cheaper

Most of the time a chat model spends writing to you, its chip is waiting for memory.

14 lessons · 146 minResearch

About this chapter

A model nobody can afford to run helps nobody. This chapter is about the price of one answer, in memory, in time and in money, and every honest way to lower it. You will round weights to four bits and see what breaks, let a big model teach a small one, cut weights away, serve a thousand customers from one base model, route each token through a few experts, let a small model draft while a big one checks, and fit a model on a phone. Then you will choose between them the only way that holds up, with a measured eval, in Arabic as well as English.

What you will be able to do

  1. 1

    Who Gets the Answer

    Work out whether a model fits a device from its weights and its bits, and name the four ways to shrink it.

    9 min
  2. 2

    The Wall Is Memory

    Explain why writing one token is limited by memory speed, and put a ceiling on tokens a second.

    10 min
  3. 3

    One Scale or Many

    Round weights to 4 bits with one scale per group, and count what the scales cost.

    11 min
  4. 4

    One Number Ruins the Row

    Show how one outlier wrecks rounding, and name two ways real methods protect the rest.

    11 min
  5. 5

    Round, Then Repair

    Round weights one at a time and cancel each error with the weights still to come.

    12 min
  6. 6

    A Teacher With a Vocabulary

    Train a student on a teacher's whole row of chances, and use heat to make the small ones teach.

    11 min
  7. 7

    Cutting Weights

    Prune a layer three ways and say which pattern a chip can actually skip.

    11 min
  8. 8

    One Base, Many Adapters

    Size an adapter, and serve many customers from one shared model.

    10 min
  9. 9

    Cheap to Run, Costly to Hold

    Say what a mixture of experts saves and what it does not, on a busy server and on a phone.

    10 min
  10. 10

    A Small Model Drafts

    Explain how a big model checks a small model's draft in one pass, and why the text does not change.

    10 min
  11. 11

    How Far Ahead to Guess

    Work out the tokens one big step yields from a draft, and pick how many tokens to draft.

    10 min
  12. 12

    The Long Conversation Bill

    Stack the levers that shrink a long conversation's cache, and say what each one gives up.

    10 min
  13. 13

    A Model in Your Pocket

    Fit a language model into a phone's share of memory, and say what the tokenizer does to an Arabic reply.

    10 min
  14. 14

    Choosing With a Measured Eval

    Pick the cheapest setting that meets a bar set in advance, measured in every language your users speak.

    11 min

Before you start

Keep going

All chapters