How AGI could go wrong

07 Mechanisms

The mechanisms behind the worst-case outcomes

Chapter 02 covers the general loss-of-control mechanism. This chapter is the register of specific, named pathways researchers actually study — because the intervention that prevents one does not necessarily prevent another.

Revised
Sources
3 cited

Why this isn't one scenario

"AI could kill everyone" compresses several structurally different failure modes into a single image. That compression is useful for a headline and useless for prevention, because the technical and policy interventions that address one mechanism often do nothing for another — a compute-governance regime that catches a rogue training run does not stop a state actor from directing an already-deployed model toward a bioweapon. Three hundred and fifty AI researchers and executives, across every major lab, signed a one-sentence statement in 2023 that treats this as a live governance priority rather than science fiction:

Summarized position

Center for AI Safetyholds that mitigating the risk of extinction from AI should be a global priority alongside pandemics and nuclear war.

Center for AI Safety, Statement on AI Risk, signed by 350+ researchers and executives
safe.ai, Statement

Two broad families

Nearly every specific mechanism below falls into one of two families. Malicious use is a human deliberately directing AI capability toward mass harm — engineered pandemics, infrastructure attacks, autonomous weapons. The AI is a more capable tool in the hands of an actor who already wanted to cause that harm. Loss of control is different in kind: the system itself ends up pursuing outcomes nobody directed, through channels like reward hacking, gradual disempowerment, or bad multi-agent equilibria. One of the field's most-cited safety researchers put the distinction plainly:

Summarized position

Yoshua Bengiowarns that a loss of control could follow from AI systems developing their own drive to survive.

Yoshua Bengio, Turing Award laureate; lead author, International AI Safety Report
AFP, via TechXplore, Secondary

A Turing Award laureate who spent decades building the field put a number on how seriously he takes the second family:

Summarized position

Geoffrey Hintonestimates a 10–20% chance that AI drives human extinction within thirty years.

Geoffrey Hinton, Turing Award laureate; former VP, Google
BBC Radio 4, reported by Forbes, Secondary

What this chapter does not do

Editorial boundary

Every page in this chapter is written at the same level its cited sources use: naming a risk category, its real-world evidence, and the mitigations researchers propose. None describes synthesis, exploit, or attack specifics — that information would not make a reader safer, and publishing it is not what informed preparation requires. See the sourcing standard for how every claim on this site is checked.

In this chapter

6pages in this chapter

Gradual disempowermentHow humans could lose meaningful control of civilization's key systems through accumulated, individually reasonable decisions — no takeover required.
Reviewed2 min
Reward hacking and specification gamingHow AI systems satisfy the literal reward they're given while defeating the designer's actual intent — a documented, growing catalog of real cases.
Reviewed2 min

Type to search the manual.

navigate open esc close