Stage 2 · Intermediate · L3
The Transformer
One architecture ate the whole field.
13 lessons · 134 minDemanding
About this chapter
A transformer block has two parts that work, attention and a feed-forward layer, and two norms and two additions that keep it steady. Stack the block and you have the models in the news. This chapter opens the block, follows the residual stream through the stack, and shows why the feed-forward layers hold most of the parameters. You will count a real model's parameters by hand, then train a tiny transformer on this device.
What you will be able to do
- 110 min
Anatomy of One Block
Name the parts of a transformer block and their order.
- 210 min
The Feed-Forward Layer
Explain the widen-and-narrow shape and count its parameters.
- 310 min
Norms and Residuals
Explain pre-norm versus post-norm and why pre-norm trains more easily.
- 410 min
The Residual Stream
Describe the stack as layers reading from and writing to one shared stream.
- 510 min
Positions in Practice
Say where rotary positions act in the block and how context length gets extended.
- 611 min
Three Shapes of Transformer
Choose an encoder, a decoder or both for a given task.
- 710 min
GPT and BERT
Compare the two training objectives and what each model is good for.
- 810 min
Where Facts Live
Summarize the evidence that feed-forward layers act as key-value memories.
- 911 min
Mixture of Experts
Explain routing and the gap between total and active parameters.
- 109 min
From Vector to Next Token
Trace the last vector through the output head to a probability for every token.
- 1110 min
Why It Beat the RNN
Explain why a transformer trains in parallel and a recurrent network cannot.
- 1211 min
Reading a Real Config
Read layers, heads, width and vocabulary, and count a model's parameters by hand.
- 1312 min
Build a Tiny Transformer
Assemble the parts into a working model and train it on this device.