The word strawberry has three r's, and a model cannot see any of them.
L1 · Words as Numbers
A language model never sees letters. It sees token ids, and the way text is cut into tokens quietly decides what the model finds easy or impossible. This chapter builds byte pair encoding merge by merge, then turns ids into embeddings, which are coordinates the model learns for meaning. You will see why arithmetic on words sometimes works, and measure the extra cost Arabic pays on most tokenizers. Many strange model failures start at this step.
- Text Has to Become Numbers
- Characters, Words, Subwords
- Byte Pair Encoding, Merge by Merge
- How Big Should the Vocabulary Be?
- Special Tokens and Chat Templates
- Embeddings Are Learned Coordinates
- Similarity and Analogies
- Meaning From Company
- Where a Token Sits
- The Weird Failures
- What Arabic Costs a Tokenizer
- Counting Tokens and Money