D3 · Backpropagation
Gradients That Fade
Explain why a gradient shrinks with depth, and measure how fast.
A thirty layer sigmoid network is not slow to train at the bottom. It does not train there at all. The gradient that arrives is smaller than the rounding error in the numbers holding it.