L2 · Attention
Cheaper Attention Variants
Compare multi-query, grouped-query and latent attention on cache size.
Look at what the cache holds: keys and values, for every head, layer and token. No queries: a query is used once and dropped. So the cost is in the filing, not the asking. How few keys and values can a layer keep and still ask many questions?