L2 · Attention
Flash Attention
Explain how never writing the full attention matrix saves memory and time.
A GPU multiplies far faster than it fetches. Its on-chip memory is over ten times faster than the main memory beside it, and about two thousand times smaller. Plain attention writes the whole score grid out and reads it back, more than once. Most of the time goes to moving, not multiplying.