Skip to content

Stage 2 · Intermediate · L1

Words as Numbers

The word strawberry has three r's, and a model cannot see any of them.

12 lessons · 117 minSteady

About this chapter

A language model never sees letters. It sees token ids, and the way text is cut into tokens decides what the model finds easy or impossible. This chapter builds byte pair encoding merge by merge, then turns ids into embeddings, which are coordinates the model learns for meaning. You will see why arithmetic on words sometimes works, and measure the extra cost Arabic pays on most tokenizers. Many strange model failures start at this step.

What you will be able to do

  1. 1

    Text Has to Become Numbers

    Explain why a model needs a fixed vocabulary of integer ids.

    8 min
  2. 2

    Characters, Words, Subwords

    Compare the three tokenization levels on vocabulary size and sequence length.

    9 min
  3. 3

    Byte Pair Encoding, Merge by Merge

    Run BPE training on a small sample text by hand and read the merge list.

    11 min
  4. 4

    How Big Should the Vocabulary Be?

    Weigh vocabulary size against sequence length, memory and rare words.

    9 min
  5. 5

    Special Tokens and Chat Templates

    Read a chat template and point to the tokens that mark roles.

    8 min
  6. 6

    Embeddings Are Learned Coordinates

    Describe what an embedding table is and how it is trained.

    11 min
  7. 7

    Similarity and Analogies

    Use cosine similarity to find neighbors and test an analogy honestly.

    10 min
  8. 8

    Meaning From Company

    Explain the skip-gram training signal in one sentence and why it works.

    10 min
  9. 9

    Where a Token Sits

    Explain why order must be injected and compare absolute with rotary positions.

    10 min
  10. 10

    The Weird Failures

    Predict which tasks break because of tokenization, including counting and math.

    11 min
  11. 11

    What Arabic Costs a Tokenizer

    Measure tokens per word for Arabic and English and explain the price gap.

    12 min
  12. 12

    Counting Tokens and Money

    Estimate the token count and cost of a real prompt before you send it.

    8 min

Before you start

Keep going

All chapters