D4 · Deep Learning
The Shortcut That Opened Up Depth
Explain why adding the input back makes a very deep stack trainable.
In 2015 a team at Microsoft Research stacked 56 plain layers and compared them with 20. The 56 layer network scored worse. Not worse on new data. Worse on the very training data it was allowed to memorize.