Skip to content

Stage 2 · Intermediate · L2

Attention

One lab cut a model's memory of a conversation by 93 percent and kept every head.

11 lessons · 115 minSteady

About this chapter

Attention is a dictionary where every key matches a little. Each token asks a question, every token answers with how well it fits, and the reply is a blend. From there, heads, the causal mask and the KV cache follow. So does the cost: the grid of scores grows with the square of the text, which is why grouped queries, latent attention and flash attention exist.

What you will be able to do

  1. 1

    A Dictionary With Fuzzy Keys

    Describe attention as a lookup where every key matches a little.

    10 min
  2. 2

    Queries, Keys and Values

    Say what each of the three projections is for, using one sentence each.

    11 min
  3. 3

    Scaled Dot-Product Attention

    Compute one attention output by hand and explain the scaling factor.

    11 min
  4. 4

    Softmax Over Positions

    Read an attention row as a probability distribution over tokens.

    10 min
  5. 5

    Many Heads, Many Questions

    Explain why splitting into heads beats one wide attention.

    11 min
  6. 6

    No Peeking Ahead

    Apply a causal mask and say why training would be useless without it.

    9 min
  7. 7

    Reading Attention Maps

    Interpret a real attention map and list what it cannot prove.

    10 min
  8. 8

    The KV Cache

    Explain what is cached during generation and how memory grows with length.

    11 min
  9. 9

    Why Long Context Is Expensive

    Compute attention cost for a given length and compare two lengths.

    10 min
  10. 10

    Cheaper Attention Variants

    Going further: compare multi-query, grouped-query and latent attention on cache size.

    12 min
  11. 11

    Flash Attention

    Going further: explain how never writing the full attention matrix saves memory and time.

    10 min

Before you start

Keep going

All chapters