What this page is, and is not
This page is about the evidence behind the numbers, not the numbers themselves. You have seen the headlines: one serious researcher says the probability of AI-driven human extinction is around 20%, another says it is effectively zero, and both cite expertise. This page explains where those estimates come from, why they spread across two orders of magnitude, and — most usefully — what a careful reader should do with a quantity no one can measure.
The survey results and models discussed here are observed publications. The interpretation of what their disagreement means is analysis. For the companion argument that these numbers describe a categorically different kind of outcome, see extinction vs. mass-casualty events; the manual’s broader treatment of the estimates lives at risk estimates and expert disagreement.
What the surveys actually say
The largest relevant survey asked thousands of machine-learning researchers — authors at major AI conferences — for their probabilities on the long-run impact of advanced AI. The 2023 wave found a median estimate of about 5% for human extinction or similarly permanent and severe disempowerment caused by future AI advances, with a mean near 16% and an interquartile range spanning roughly 0% to 19% (Grace et al., 2024). Read that distribution, not just the median: a mean four times the median means a long right tail, with a substantial minority of experts above 20% and some near 50%, while another substantial minority sits essentially at zero.
The same paper documents that the number moves with the wording. When the question specified extinction caused by humans’ inability to control advanced AI, the median doubled to about 10%. When a 2022 wave asked about extinction from AI in general, the median was 5%; the earlier AI Impacts survey rounds show the same sensitivity to framing (2022 Expert Survey on Progress in AI). A quantity that shifts by a factor of two when the question is rephrased is telling you something about the respondents’ uncertainty — and, frankly, about the looseness of the concept being measured. None of this makes the surveys worthless. It makes them raw material, not verdicts.
Model-based estimates versus gut judgments
A second family of estimates does not ask anyone’s intuition; it builds an explicit argument and assigns probabilities to each step. The most rigorous example is Joseph Carlsmith’s structured analysis of whether power-seeking AI poses an existential risk, which decomposes the claim into six conjuncts — systems will become much more capable, they will seek power, humans will fail to correct this, and so on — and defends a probability on each. His bottom line lands around 10% for AI-caused catastrophe this century, with wide error bars he is careful to display (Carlsmith, 2022). Philosophers working in the same tradition reach different numbers by disputing individual conjuncts: Toby Ord, in The Precipice (2020), argues for an overall existential risk from unaligned AI of roughly one in six this century, driven by a more pessimistic reading of the alignment problem.
These models share a virtue the surveys lack: every step is visible and contestable. They share two vices. Every step is also contestable, which is why reasonable readers reverse the sign of individual conjuncts; and the probabilities attached to each step are still, in the end, judgments. A gut judgment with a diagram is still a gut judgment. The disagreement between a 10% structured model and a 0.1% intuition is not a measurement conflict. It is a conflict about which reference classes and which failure modes deserve weight.
Why reference classes pull in different directions
Underneath every estimate sits an analogy, and the analogies genuinely conflict:
- Biological extinction. Species go extinct constantly; the background rate is roughly one to five species per year in recent decades. But those are species with small populations and narrow ranges. The nearest human analogue, the Toba supereruption bottleneck, is contested and predates civilization entirely. This base rate pulls toward “extinction is a normal outcome” while telling you almost nothing about a global technological species.
- Nuclear near-misses. The Cuban Missile Crisis, the 1983 Stanislav Petrov incident, and the 1995 Norwegian rocket scare show systems under maximum tension producing near-failures and then not failing. Optimists read a long string of saved-by-luck events; pessimists read the same string as proof the luck runs out.
- Technology accidents. Airliners, reactors, and spacecraft fail at measurable rates and then get safer through hard-learned redesign. Optimists extrapolate the redesign; pessimists note that aviation did not have to succeed only once.
Each analogy is honest. Each encodes a prior, not a fact. Estimates diverge because the priors diverge, and the priors diverge because there is no direct evidence — no completed run of the experiment — to discipline any of them.
What the disagreement itself tells you
Here is the analytical point most coverage misses: the variance is the finding. If extinction probability were easy to estimate, thousands of informed experts would not produce estimates from near zero to near coin-flip. The spread is evidence about the state of knowledge, in the same way that wildly divergent weather models are evidence that the storm is genuinely unpredictable, not that the average forecast must be right. Averaging the experts into a tidy middle figure — say, 5% — creates the feeling of precision while inheriting none of the information. The 2026 International AI Safety Report, the most comprehensive synthesis attempt to date, is candid that severe-risk evidence is thin, contested, and concentrated in a handful of laboratories (International AI Safety Report 2026).
One more fact the distribution encodes: disagreement is not random. In the surveys, researchers who expect faster capability progress assign higher catastrophic-risk probabilities, and researchers at frontier laboratories assign lower ones than independent academics. Some of that is evidence about who sees what. Some of it is incentive. Readers should hold both explanations.
How to actually use these numbers
The practical conclusion is not that the numbers are useless; it is that they must be used with the right tool. Expected-value bookkeeping — multiply a precise probability by a payoff and rank the results — fails here, because the input is a two-orders-of-magnitude-wide judgment and small errors in the input silently dominate the output. Decision-making under deep uncertainty works differently: you ask which actions remain worth taking across the whole plausible range, and you prefer actions that are robust rather than optimal.
Applied to AI extinction risk, that logic is clarifying. Prevention work — capability evaluations, deployment safeguards, compute governance — looks worthwhile at 20%, at 5%, and at 1%, because the cost of the work is modest and the outcome it guards is categorically different from anything priced in ordinary policy. Dismissal looks defensible only near zero. So the surveys, despite their looseness, do real work: they tell you that near-zero is a contested position, not a consensus. For how the underlying timeline expectations drive the disagreement, see timeline forecasts. You do not need to know the probability of the fire to justify buying the extinguisher. You need to know that serious people, looking at the same evidence, keep smelling smoke.