One lab cut a model's memory of a conversation by 93 percent and kept every head.
L2 · Attention
Attention is a dictionary where every key matches a little. Each token asks a question, every token answers with how well it fits, and the reply is a blend. From there, heads, the causal mask and the KV cache follow. So does the cost: the grid of scores grows with the square of the text, which is why grouped queries, latent attention and flash attention exist.
- A Dictionary With Fuzzy Keys
- Queries, Keys and Values
- Scaled Dot-Product Attention
- Softmax Over Positions
- Many Heads, Many Questions
- No Peeking Ahead
- Reading Attention Maps
- The KV Cache
- Why Long Context Is Expensive
- Cheaper Attention Variants
- Flash Attention