Skip to content

Stage 2 · Intermediate · L3

The Transformer

One architecture ate the whole field.

13 lessons · 134 minDemanding

About this chapter

A transformer block has two parts that work, attention and a feed-forward layer, and two norms and two additions that keep it steady. Stack the block and you have the models in the news. This chapter opens the block, follows the residual stream through the stack, and shows why the feed-forward layers hold most of the parameters. You will count a real model's parameters by hand, then train a tiny transformer on this device.

What you will be able to do

  1. 1

    Anatomy of One Block

    Name the parts of a transformer block and their order.

    10 min
  2. 2

    The Feed-Forward Layer

    Explain the widen-and-narrow shape and count its parameters.

    10 min
  3. 3

    Norms and Residuals

    Explain pre-norm versus post-norm and why pre-norm trains more easily.

    10 min
  4. 4

    The Residual Stream

    Describe the stack as layers reading from and writing to one shared stream.

    10 min
  5. 5

    Positions in Practice

    Say where rotary positions act in the block and how context length gets extended.

    10 min
  6. 6

    Three Shapes of Transformer

    Choose an encoder, a decoder or both for a given task.

    11 min
  7. 7

    GPT and BERT

    Compare the two training objectives and what each model is good for.

    10 min
  8. 8

    Where Facts Live

    Summarize the evidence that feed-forward layers act as key-value memories.

    10 min
  9. 9

    Mixture of Experts

    Explain routing and the gap between total and active parameters.

    11 min
  10. 10

    From Vector to Next Token

    Trace the last vector through the output head to a probability for every token.

    9 min
  11. 11

    Why It Beat the RNN

    Explain why a transformer trains in parallel and a recurrent network cannot.

    10 min
  12. 12

    Reading a Real Config

    Read layers, heads, width and vocabulary, and count a model's parameters by hand.

    11 min
  13. 13

    Build a Tiny Transformer

    Assemble the parts into a working model and train it on this device.

    12 min

Before you start

Keep going

All chapters