D2 · Gradient Descent
A Different Step for Every Weight
Explain why one shared learning rate is a bad fit for real models, and how RMSProp gives each weight its own scale.
So far, every weight in a model gets the same learning rate. Picture a model that prices apartments in Riyadh. One input is the area, around 200 square meters. Another is the number of rooms, around 4.