How AGI could go wrong

Gradual disempowerment

How humans could lose meaningful control of civilization's key systems through accumulated, individually reasonable decisions — no takeover required.

Status
Reviewed
Revised
Sources
2 cited
Reading
2 min

A different shape of losing control

Most catastrophic AI scenarios that reach the public imagination involve a dramatic moment — a system breaks free, seizes infrastructure, or otherwise announces itself. Gradual disempowerment, a concept formalized in a 2025 paper by researchers including Jan Kulveit, David Krueger, and David Duvenaud, describes a mechanism with no such moment at all: humans could lose meaningful control over civilization’s key systems — the economy, the state, and culture — through a long accumulation of individually reasonable decisions, none of which resembles a takeover. This complements rather than replaces the more abrupt loss-of-control mechanism described elsewhere in this manual; the two are different failure shapes, not competing predictions.

The mechanism, in outline

The argument runs through institutions rather than through any single rogue system. Companies replace human judgment with AI decision-making because it’s cheaper and faster, which is individually rational for each firm; governments come to rely on AI-driven analysis and administration for the same reason; cultural and informational systems increasingly optimize for AI-mediated attention rather than human preference. Each shift weakens a different lever humans have historically used to keep institutions responsive to their interests — consumer choice, voting, cultural feedback — not because anyone intends to disempower people, but because AI adoption at each individual decision point is locally sensible and only produces the disempowering pattern in aggregate. Formal human authority — voting rights, legal ownership, nominal command — can remain fully intact throughout, even as the actual capacity to exercise it erodes.

Why this doesn’t require a villain

The idea is close in spirit to an influential 2019 essay arguing that AI catastrophe could emerge from a slow-rolling failure, in which optimizing for measurable proxies of human approval gradually diverges from what people actually want, without any single decision anyone could point to as the mistake. Both accounts share the same unsettling feature — each step along the way looks like an improvement, or at worst a defensible tradeoff, by the standards available at the time it’s made.

How it connects to the rest of the risk picture

Gradual disempowerment is also a power-concentration story, just distributed differently: rather than one company or state accumulating dominant control (see power concentration), authority can migrate away from human institutions generally and toward optimization processes that no one — including their nominal owners — fully steers. That reframes the policy question: the safeguard isn’t only “keep AI out of the wrong hands,” but “keep enough real human judgment load-bearing in economic and political systems that it doesn’t quietly become vestigial.”

Sources

Summarized position

Jan Kulveit, Raymond Douglas, Nora Ammann, Deger Turan, David Krueger, and David Duvenaudargues humans could lose control of civilization's key systems through accumulated, individually rational decisions rather than a takeover.

Jan Kulveit, Raymond Douglas, Nora Ammann, Deger Turan, David Krueger, and David Duvenaud, Authors, "Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development"
arXiv, Primary
Summarized position

Paul Christianoargues AI catastrophe more plausibly emerges from many systems optimizing measurable proxies and developing influence-seeking tendencies than from one rogue superintelligence.

Paul Christiano, AI alignment researcher; former OpenAI
"What Failure Looks Like" (AI Alignment Forum), Primary

Full register of everything this manual cites: source index.

Type to search the manual.

navigate open esc close