D3 · Backpropagation
Gradients That Explode, Units That Die
Spot exploding gradients and dead ReLUs from the numbers you log.
Picture a model training cleanly for two hours. Then one batch gives a gradient norm near 1,200,000, and every loss after it prints NaN. The loss curve warned of nothing. The gradient norm showed it all.