R3 · Open Problems
Alignment as an Open Problem
Explain scalable oversight and why current methods may not extend.
Today's method for teaching a model to behave rests on one assumption. A person can look at two answers and say which is better. Every part of the method needs that to be true.