D6 · Sequences & Memory
The Same Problem, Now in Time
Show why long dependencies vanish and where the signal is lost.
A recurrent network can carry a fact forward for a thousand steps. Going forward is easy. The trouble is the way back. To learn that step 1 mattered, a gradient must travel there from the loss, and it fades. So the model does not forget. It is never told that remembering would have paid.