Oversight and governance

Scalable oversight: supervising more capable systems

The technical approaches researchers are testing to let human overseers meaningfully supervise AI systems more capable than themselves.

Status
Working draft
Revised
Sources
None yet
Reading
1 min

The problem this solves

Standard oversight assumes the supervisor understands the task at least as well as the system being supervised. That assumption breaks down for a system operating beyond human-expert level in a domain — how do you verify work you can’t independently check?

Leading approaches

  • Debate — two AI systems argue opposing sides of a question in front of a human or weaker AI judge, on the theory that it’s easier to spot a flaw in an opposing argument than to independently verify a complex claim from scratch.
  • Recursive reward modeling — using AI assistance to help humans evaluate AI outputs on tasks too complex to judge unaided, building up supervisory capability in layers.
  • Weak-to-strong generalization — studying whether a weaker model’s supervision signal can still reliably steer a stronger model’s behavior in the right direction, even though the weaker model can’t fully verify the stronger one’s reasoning.

Where the research stands

These techniques are active, published research directions with encouraging but early results in controlled settings — none has been demonstrated at the scale or stakes a genuinely superhuman system would require. This is one of the areas where the gap between current research maturity and the capability trajectory (see timeline forecasts) is most concerning to safety researchers.

Why this can’t be solved by regulation alone

Scalable oversight is fundamentally a technical problem — no law can make an unsolved verification problem solved. Governance (see international treaties and summits) can buy time and set incentives for labs to invest in this research; it can’t substitute for the research succeeding.

Status: working draft

Real analysis at working-draft depth. Treat specifics as provisional until sourced. This page has been revised as recently as any other — see the revision log — but its specifics are not yet backed by citations on the page itself.

Type to search the manual.

navigate open esc close