L4 · Training a Language Model
DPO, and Doing Less
Compare DPO with RLHF on moving parts, and say where written rules replace human labels.
In 2023 a Stanford team showed that the reward model can be skipped. The language model already gives every reply a chance. Train on the preferences directly, and the separate scorer disappears.