R2 · Landmark Papers, Rebuilt
InstructGPT (2022)
Fit a reward model from comparisons, and say what human preference data bought.
A next-word model is trained to continue text. Ask it a question and a plausible continuation is another question. The model was not being unhelpful. It was doing exactly the job it was trained on.