If prevention fails

Why expert extinction estimates diverge so widely

What expert surveys, structured models, and historical base rates actually say about AI extinction probability, and how to use numbers this uncertain.

Written by
Dwight Ringdahl
Status
Reviewed
Revised
Sources
4 cited
Reading
5 min

What this page is, and is not

This page is about the evidence behind the numbers, not the numbers themselves. You have seen the headlines: one serious researcher says the probability of AI-driven human extinction is around 20%, another says it is effectively zero, and both cite expertise. This page explains where those estimates come from, why they spread across two orders of magnitude, and — most usefully — what a careful reader should do with a quantity no one can measure.

The survey results and models discussed here are observed publications. The interpretation of what their disagreement means is analysis. For the companion argument that these numbers describe a categorically different kind of outcome, see extinction vs. mass-casualty events; the manual’s broader treatment of the estimates lives at risk estimates and expert disagreement.

What the surveys actually say

The largest relevant survey asked thousands of machine-learning researchers — authors at major AI conferences — for their probabilities on the long-run impact of advanced AI. The 2023 wave found a median estimate of about 5% for human extinction or similarly permanent and severe disempowerment caused by future AI advances, with a mean near 16% and an interquartile range spanning roughly 0% to 19% (Grace et al., 2024). Read that distribution, not just the median: a mean four times the median means a long right tail, with a substantial minority of experts above 20% and some near 50%, while another substantial minority sits essentially at zero.

The same paper documents that the number moves with the wording. When the question specified extinction caused by humans’ inability to control advanced AI, the median doubled to about 10%. When a 2022 wave asked about extinction from AI in general, the median was 5%; the earlier AI Impacts survey rounds show the same sensitivity to framing (2022 Expert Survey on Progress in AI). A quantity that shifts by a factor of two when the question is rephrased is telling you something about the respondents’ uncertainty — and, frankly, about the looseness of the concept being measured. None of this makes the surveys worthless. It makes them raw material, not verdicts.

Model-based estimates versus gut judgments

A second family of estimates does not ask anyone’s intuition; it builds an explicit argument and assigns probabilities to each step. The most rigorous example is Joseph Carlsmith’s structured analysis of whether power-seeking AI poses an existential risk, which decomposes the claim into six conjuncts — systems will become much more capable, they will seek power, humans will fail to correct this, and so on — and defends a probability on each. His bottom line lands around 10% for AI-caused catastrophe this century, with wide error bars he is careful to display (Carlsmith, 2022). Philosophers working in the same tradition reach different numbers by disputing individual conjuncts: Toby Ord, in The Precipice (2020), argues for an overall existential risk from unaligned AI of roughly one in six this century, driven by a more pessimistic reading of the alignment problem.

These models share a virtue the surveys lack: every step is visible and contestable. They share two vices. Every step is also contestable, which is why reasonable readers reverse the sign of individual conjuncts; and the probabilities attached to each step are still, in the end, judgments. A gut judgment with a diagram is still a gut judgment. The disagreement between a 10% structured model and a 0.1% intuition is not a measurement conflict. It is a conflict about which reference classes and which failure modes deserve weight.

Why reference classes pull in different directions

Underneath every estimate sits an analogy, and the analogies genuinely conflict:

  • Biological extinction. Species go extinct constantly; the background rate is roughly one to five species per year in recent decades. But those are species with small populations and narrow ranges. The nearest human analogue, the Toba supereruption bottleneck, is contested and predates civilization entirely. This base rate pulls toward “extinction is a normal outcome” while telling you almost nothing about a global technological species.
  • Nuclear near-misses. The Cuban Missile Crisis, the 1983 Stanislav Petrov incident, and the 1995 Norwegian rocket scare show systems under maximum tension producing near-failures and then not failing. Optimists read a long string of saved-by-luck events; pessimists read the same string as proof the luck runs out.
  • Technology accidents. Airliners, reactors, and spacecraft fail at measurable rates and then get safer through hard-learned redesign. Optimists extrapolate the redesign; pessimists note that aviation did not have to succeed only once.

Each analogy is honest. Each encodes a prior, not a fact. Estimates diverge because the priors diverge, and the priors diverge because there is no direct evidence — no completed run of the experiment — to discipline any of them.

What the disagreement itself tells you

Here is the analytical point most coverage misses: the variance is the finding. If extinction probability were easy to estimate, thousands of informed experts would not produce estimates from near zero to near coin-flip. The spread is evidence about the state of knowledge, in the same way that wildly divergent weather models are evidence that the storm is genuinely unpredictable, not that the average forecast must be right. Averaging the experts into a tidy middle figure — say, 5% — creates the feeling of precision while inheriting none of the information. The 2026 International AI Safety Report, the most comprehensive synthesis attempt to date, is candid that severe-risk evidence is thin, contested, and concentrated in a handful of laboratories (International AI Safety Report 2026).

One more fact the distribution encodes: disagreement is not random. In the surveys, researchers who expect faster capability progress assign higher catastrophic-risk probabilities, and researchers at frontier laboratories assign lower ones than independent academics. Some of that is evidence about who sees what. Some of it is incentive. Readers should hold both explanations.

How to actually use these numbers

The practical conclusion is not that the numbers are useless; it is that they must be used with the right tool. Expected-value bookkeeping — multiply a precise probability by a payoff and rank the results — fails here, because the input is a two-orders-of-magnitude-wide judgment and small errors in the input silently dominate the output. Decision-making under deep uncertainty works differently: you ask which actions remain worth taking across the whole plausible range, and you prefer actions that are robust rather than optimal.

Applied to AI extinction risk, that logic is clarifying. Prevention work — capability evaluations, deployment safeguards, compute governance — looks worthwhile at 20%, at 5%, and at 1%, because the cost of the work is modest and the outcome it guards is categorically different from anything priced in ordinary policy. Dismissal looks defensible only near zero. So the surveys, despite their looseness, do real work: they tell you that near-zero is a contested position, not a consensus. For how the underlying timeline expectations drive the disagreement, see timeline forecasts. You do not need to know the probability of the fire to justify buying the extinguisher. You need to know that serious people, looking at the same evidence, keep smelling smoke.

References

Summarized position

AI Impacts found the same sensitivity to question wording in its 2022 round of AI-researcher surveys as the larger 2023 wave, with a median 5% estimate for extinction from AI in general.

AI Impacts, 2022 Expert Survey on Progress in AI
AI Impacts, Report
Summarized position

Joseph Carlsmith decomposes the case for AI-caused existential catastrophe into six conjuncts and defends an overall probability of roughly 10% for this century, with wide displayed uncertainty.

Joseph Carlsmith, Author, "Is Power-Seeking AI an Existential Risk?"
arXiv, Primary
Summarized position

International AI Safety Report found that severe AI risk pathways — including loss of control, large-scale disruption, and deliberate misuse — are plausible enough to warrant mitigation work now, despite a still-thin evidence base.

International AI Safety Report, 2026 edition, chaired by Yoshua Bengio
internationalaisafetyreport.org, Report
  1. Grace et al., 2024 arxiv.org

The source index also tracks the manual's recurring core sources and expert positions.

Type to search the manual.

navigate open esc close