D4 · Deep Learning
Where the Weights Start Matters
Choose Xavier or He initialization and say what each keeps constant.
Give every weight in a layer the same starting number, and its units compute the same thing. They get the same gradient and change by the same amount, forever. The layer has three units on paper and one in practice.