Stage 2 · Intermediate · L2
Attention
One lab cut a model's memory of a conversation by 93 percent and kept every head.
11 lessons · 115 minSteady
About this chapter
Attention is a dictionary where every key matches a little. Each token asks a question, every token answers with how well it fits, and the reply is a blend. From there, heads, the causal mask and the KV cache follow. So does the cost: the grid of scores grows with the square of the text, which is why grouped queries, latent attention and flash attention exist.
What you will be able to do
- 110 min
A Dictionary With Fuzzy Keys
Describe attention as a lookup where every key matches a little.
- 211 min
Queries, Keys and Values
Say what each of the three projections is for, using one sentence each.
- 311 min
Scaled Dot-Product Attention
Compute one attention output by hand and explain the scaling factor.
- 410 min
Softmax Over Positions
Read an attention row as a probability distribution over tokens.
- 511 min
Many Heads, Many Questions
Explain why splitting into heads beats one wide attention.
- 69 min
No Peeking Ahead
Apply a causal mask and say why training would be useless without it.
- 710 min
Reading Attention Maps
Interpret a real attention map and list what it cannot prove.
- 811 min
The KV Cache
Explain what is cached during generation and how memory grows with length.
- 910 min
Why Long Context Is Expensive
Compute attention cost for a given length and compare two lengths.
- 1012 min
Cheaper Attention Variants
Going further: compare multi-query, grouped-query and latent attention on cache size.
- 1110 min
Flash Attention
Going further: explain how never writing the full attention matrix saves memory and time.