Scalable oversight of superhuman models
supervising models you cannot fully evaluate.
humans cannot judge outputs beyond their expertise, so feedback degrades; techniques like AI debate or recursive reward modeling decompose judgment.
WHY IT IS HARD Scalable oversight is the problem of supervising and giving training feedback to models whose capabilities exceed those of their human supervisors. When a model can write code, prove theorems, or summarize research better than the rater, the human cannot reliably tell a correct answer from a plausible but subtly wrong one. Human feedback then becomes noisy or systematically biased, and optimizing against it can reward persuasive errors rather than truth.
A GOOD ANSWER COVERS The core gap is verification, not effort: adding more raters does not help if none of them can judge the domain. Proposed techniques decompose or assist the judgment. In debate, two model instances argue opposing positions on a question and a human judges which argument is more truthful; the bet is that exposing a lie is easier than producing it. In recursive reward modeling, you train assistant models to help humans evaluate harder tasks, bootstrapping oversight of capability level n using trusted models at level n minus one. Task decomposition and process-based supervision also reduce the load on any single human judgment.
WHEN IT MATTERS This becomes critical exactly when models approach or surpass expert humans, because that is when ordinary RLHF stops giving a reliable signal.
COMMON WRONG ANSWERS Saying just hire domain experts, which fails once models exceed the best experts, or assuming larger preference datasets resolve a fundamental verification problem.
ONE CONCRETE EXAMPLE Ask a model a subtle question whose answer a non-expert cannot verify, such as whether a long proof is valid. In a debate setup, a second model points to the exact flawed step, and the human, unable to check the whole proof, can nonetheless evaluate that focused claim and reward the honest model.
Get five bites like this every day.
Tezvyn delivers a daily feed of 60-second tech bites with quizzes to lock in what you learn.