Stage 2 · Intermediate · D6
Sequences & Memory
Before attention, a model had to squeeze a whole paragraph into one vector.
11 lessons · 107 minSteady
About this chapter
Order carries meaning, so a model that reads a sentence needs a memory. Recurrent networks gave it one: a hidden state passed from step to step. This chapter unrolls that loop, trains it through time, and shows the exact point where gradients die on long inputs. Gates fixed part of it, and a single bottleneck vector in translation exposed the rest. Attention was invented to close that gap.
What you will be able to do
- 18 min
Data Where Order Matters
Recognize sequence problems and say why a fixed input size fails them.
- 29 min
A Vector That Remembers
Describe the hidden state as a running summary of everything seen so far.
- 310 min
Unrolling the Loop
Draw an RNN unrolled over time and point to the shared weights.
- 411 min
Backprop Through Time
Run backprop across time steps and explain truncation.
- 510 min
The Same Problem, Now in Time
Show why long dependencies vanish and where the signal is lost.
- 612 min
Gates That Choose What to Keep
Walk one time step through an LSTM's three gates and cell state.
- 78 min
GRU: Fewer Gates, Similar Result
Compare GRU to LSTM on parameters, speed and typical accuracy.
- 89 min
Training vs Generating
Explain teacher forcing and the exposure gap it creates at generation time.
- 911 min
A Character-Level Language Model
Train a tiny character model on this device and sample text from it.
- 1010 min
One Vector for a Whole Sentence
Explain the encoder-decoder bottleneck and measure how it hurts long inputs.
- 119 min
The Case for Attention
State the three problems attention solves that gates could not.