Skip to content

Stage 2 · Intermediate · D6

Sequences & Memory

Before attention, a model had to squeeze a whole paragraph into one vector.

11 lessons · 107 minSteady

About this chapter

Order carries meaning, so a model that reads a sentence needs a memory. Recurrent networks gave it one: a hidden state passed from step to step. This chapter unrolls that loop, trains it through time, and shows the exact point where gradients die on long inputs. Gates fixed part of it, and a single bottleneck vector in translation exposed the rest. Attention was invented to close that gap.

What you will be able to do

  1. 1

    Data Where Order Matters

    Recognize sequence problems and say why a fixed input size fails them.

    8 min
  2. 2

    A Vector That Remembers

    Describe the hidden state as a running summary of everything seen so far.

    9 min
  3. 3

    Unrolling the Loop

    Draw an RNN unrolled over time and point to the shared weights.

    10 min
  4. 4

    Backprop Through Time

    Run backprop across time steps and explain truncation.

    11 min
  5. 5

    The Same Problem, Now in Time

    Show why long dependencies vanish and where the signal is lost.

    10 min
  6. 6

    Gates That Choose What to Keep

    Walk one time step through an LSTM's three gates and cell state.

    12 min
  7. 7

    GRU: Fewer Gates, Similar Result

    Compare GRU to LSTM on parameters, speed and typical accuracy.

    8 min
  8. 8

    Training vs Generating

    Explain teacher forcing and the exposure gap it creates at generation time.

    9 min
  9. 9

    A Character-Level Language Model

    Train a tiny character model on this device and sample text from it.

    11 min
  10. 10

    One Vector for a Whole Sentence

    Explain the encoder-decoder bottleneck and measure how it hurts long inputs.

    10 min
  11. 11

    The Case for Attention

    State the three problems attention solves that gates could not.

    9 min

Before you start

Keep going

All chapters