L3 · The Transformer
Norms and Residuals
Explain pre-norm versus post-norm and why pre-norm trains more easily.
The 2017 transformer tidied its numbers after each addition. Training it took care: the learning rate had to start near zero and creep up. GPT-2 moved each norm to the front of its block. Same parts, one line moved, and deep stacks became much easier to train.