One architecture ate the whole field.
L3 · The Transformer
A transformer block has two parts that work, attention and a feed-forward layer, and two norms and two additions that keep it steady. Stack the block and you have the models in the news. This chapter opens the block, follows the residual stream through the stack, and shows why the feed-forward layers hold most of the parameters. You will count a real model's parameters by hand, then train a tiny transformer on this device.
- Anatomy of One Block
- The Feed-Forward Layer
- Norms and Residuals
- The Residual Stream
- Positions in Practice
- Three Shapes of Transformer
- GPT and BERT
- Where Facts Live
- Mixture of Experts
- From Vector to Next Token
- Why It Beat the RNN
- Reading a Real Config
- Build a Tiny Transformer