L2 · Attention
Queries, Keys and Values
Say what each of the three projections is for, using one sentence each.
Every token arrives as one vector. Attention makes three versions of it, each through its own weight matrix. The query version asks. The key version gets compared. The value version gets handed over. Nobody writes those matrices by hand. Training learns them, like every other weight.