L4 · Training a Language Model
Learning From Preferences
Describe how two replies and a human choice become a score, and why a leash is needed.
Few people can write the perfect answer to a hard question. Almost anyone can read two answers and say which is better. That gap is why this stage exists.