Stage 2 · Intermediate · L1
Words as Numbers
The word strawberry has three r's, and a model cannot see any of them.
12 lessons · 117 minSteady
About this chapter
A language model never sees letters. It sees token ids, and the way text is cut into tokens decides what the model finds easy or impossible. This chapter builds byte pair encoding merge by merge, then turns ids into embeddings, which are coordinates the model learns for meaning. You will see why arithmetic on words sometimes works, and measure the extra cost Arabic pays on most tokenizers. Many strange model failures start at this step.
What you will be able to do
- 18 min
Text Has to Become Numbers
Explain why a model needs a fixed vocabulary of integer ids.
- 29 min
Characters, Words, Subwords
Compare the three tokenization levels on vocabulary size and sequence length.
- 311 min
Byte Pair Encoding, Merge by Merge
Run BPE training on a small sample text by hand and read the merge list.
- 49 min
How Big Should the Vocabulary Be?
Weigh vocabulary size against sequence length, memory and rare words.
- 58 min
Special Tokens and Chat Templates
Read a chat template and point to the tokens that mark roles.
- 611 min
Embeddings Are Learned Coordinates
Describe what an embedding table is and how it is trained.
- 710 min
Similarity and Analogies
Use cosine similarity to find neighbors and test an analogy honestly.
- 810 min
Meaning From Company
Explain the skip-gram training signal in one sentence and why it works.
- 910 min
Where a Token Sits
Explain why order must be injected and compare absolute with rotary positions.
- 1011 min
The Weird Failures
Predict which tasks break because of tokenization, including counting and math.
- 1112 min
What Arabic Costs a Tokenizer
Measure tokens per word for Arabic and English and explain the price gap.
- 128 min
Counting Tokens and Money
Estimate the token count and cost of a real prompt before you send it.