Start with the kind of claim being made
“AI can do this,” “AGI will arrive soon,” and “this system could cause catastrophe” are not the same claim. They need different evidence. A polished demonstration may establish that a model succeeded once under favorable conditions. It does not establish reliability, economic value, or safe autonomous operation. A forecast expresses a judgment about the future. It is not a measurement of a present capability.
This matters because AGI has no universally accepted test. Definitions emphasize different combinations of breadth, performance, autonomy, learning, physical competence, and economic impact. Rather than asking whether one headline proves AGI, readers should ask what was measured, under which conditions, and how far the conclusion travels beyond the evidence.
The ladder below is not a rigid ranking in which every item on one level defeats everything below it. A reproducible laboratory evaluation may be stronger evidence for a narrow mechanism than noisy deployment data. Its purpose is to expose the inferential steps that promotional and alarmist claims often hide.
Level one: anecdote and curated demonstration
Anecdotes reveal possibilities and failure modes. A screen recording of an agent building an application shows that at least one run produced the displayed result, assuming the demonstration is authentic. A surprising hallucination shows that a failure can occur. Neither tells us how often the behavior appears.
Ask whether the presenter selected the best run, edited pauses or failures, used hidden human assistance, or chose a familiar task. Look for the full prompt, tools, model version, sampling settings, number of attempts, and failure cases. Without them, treat the result as a lead worth investigating—not a rate.
Demonstrations remain useful. New abilities often appear in messy form before formal measurement exists. The mistake is turning “can happen” into “usually happens,” then into “will transform the economy.”
Level two: benchmark or controlled evaluation
An evaluation defines tasks, conditions, and a scoring rule. It can compare systems more systematically than an anecdote. Strong evaluations use held-out tasks, prevent contamination, report uncertainty, document scaffolding and human assistance, and test relevant failure modes. Independent replication increases confidence.
Benchmark scores can still mislead. Training data may contain similar questions. Teams may optimize to a public test. Multiple-choice knowledge does not establish skill in an unfamiliar workplace. Averages conceal catastrophic tail failures. A system that succeeds 90 percent of the time may be unusable in a 30-step workflow because errors compound.
The International AI Safety Report 2026 uses the term “evaluation gap” for the problem that results from controlled pre-deployment tests do not reliably predict real-world performance. It also describes current capability profiles as “jagged”: performance varies substantially across tasks and contexts, and systems can fail on apparently simple work despite succeeding on difficult evaluations. These are findings about the limits of measurement, not proof that every benchmark is useless. A credible article should report both the score and what the test leaves out.
Level three: adversarial, causal, and uplift studies
Risk claims need more than a standard benchmark. Red teams look for ways to elicit prohibited or dangerous behavior. Causal experiments change one feature and test whether outcomes change. Uplift studies compare what people can accomplish with AI against a meaningful baseline, such as the internet alone.
These studies answer different questions. A red team finding shows accessibility under the tested attack, not population-wide incidence. A capability evaluation may show that a model can produce a harmful plan without showing that a real actor can execute it. An uplift study may be more relevant to misuse, but its participants, time limit, baseline, outcome measure, and safeguards determine what it means.
Small samples produce wide uncertainty. Expert participants may understate novice uplift but better represent sophisticated attackers. Proxy tasks avoid creating actual harm but may omit real-world bottlenecks. Results age quickly as models and controls change. Good reporting names those limits instead of compressing them into “AI enables” or “AI does not increase” risk.
Level four: field and deployment evidence
Real-world data can reveal adoption, productivity, incidents, abuse, and organizational adaptation. Randomized field experiments and carefully designed quasi-experiments can estimate causal effects. Incident databases and complaint data can show recurring harm. Audits can detect disparities that a developer’s laboratory did not test.
Deployment evidence also has blind spots. Companies observe more than they disclose. Users choose whether and how to adopt a tool, making simple comparisons biased. Reported incidents are not the same as all incidents. A productivity gain in customer support does not automatically generalize to research, management, medicine, or robotics.
Readers should look for the population, setting, comparison group, time period, outcome, and conflicts of interest. “Workers completed more tasks” differs from “the business created more value,” and both differ from “employment fell.” Present labor evidence about generative AI should not be relabeled evidence about hypothetical AGI.
Level five: trend extrapolation and forecast
Forecasts combine measurements with assumptions. Compute trends, algorithmic progress, investment, benchmark performance, and agent reliability can inform a forecast, but none uniquely determines an AGI date. Constraints may ease, bind, or shift. Definitions may move.
METR’s time-horizon research estimates the human-expert completion time of tasks at which a tested agent is predicted to succeed with a stated reliability, commonly 50 or 80 percent. Its current suite is made up primarily of self-contained, well-specified software-engineering, machine-learning, and cybersecurity tasks with clear scoring rules. The measure is useful evidence about a defined agent setup and task distribution. It is not the time the AI runs autonomously, a measure of all economically valuable work, or a timer counting down to AGI. METR also cautions that measurements above 16 hours are unreliable with its current task suite.
Evaluate a forecast by its operational definition, date, probability, model, base rate, and track record. A calibrated forecast says, in effect, “outcomes assigned 30 percent should occur about 30 percent of the time.” A confident date without resolution criteria cannot be scored. Expert surveys capture informed beliefs but do not convert disagreement into experimental evidence.
Scenarios serve another purpose. They explore conditional pathways: if capabilities, access, autonomy, and institutional response take specified values, what follows? A scenario can expose vulnerabilities without being the most likely future. Label it accordingly.
Level six: mechanism and theoretical argument
Some severe AGI risks cannot be demonstrated directly without the relevant systems—and should not be created merely to obtain proof. Researchers instead propose mechanisms such as reward hacking, instrumental resource seeking, competitive races, or gradual delegation of authority. Formal models can clarify whether a conclusion follows from assumptions.
The central question is not whether the argument sounds vivid. It is which premises have empirical support. Does the system have a persistent objective? Can it plan across the required horizon? Does it have access to tools, money, infrastructure, or copies of itself? Will monitoring fail? Can institutions respond? Each link can change the probability of the whole chain.
Absence of direct catastrophe data is not evidence of zero risk. Nor does a logically possible pathway establish a large probability. Decisions under uncertainty should consider consequence, reversibility, warning time, and the cost and effectiveness of safeguards. That is a risk-management judgment, not a discovery that theory has become observation.
Use an evidence card for consequential claims
For any major statement, record six fields:
- Claim: the narrow proposition actually asserted.
- Status: observed result, inference, forecast, scenario, or recommendation.
- Source: primary study, official data, synthesis, company report, media, or opinion.
- Scope: model, version, task, population, jurisdiction, and date.
- Uncertainty: sample size, confidence interval, replication, missing data, and plausible alternatives.
- Update trigger: evidence that would strengthen, weaken, or retire the claim.
NIST’s voluntary Generative AI Profile is a cross-sector companion to the AI Risk Management Framework, not an AGI test or certification. It organizes suggested actions under the framework’s Govern, Map, Measure, and Manage functions and treats risk management as a lifecycle activity. The broader NIST AI RMF measurement guidance calls for documented testing, measures of uncertainty, deployment-relevant measurement, monitoring, and independent review where it can reduce internal bias. The exact template matters less than making the evidentiary chain auditable.
Match confidence to evidence
Reliable AGI coverage does not require false neutrality. Some facts are well established: models are improving on many evaluations, real harms already occur, capabilities remain uneven, and institutions have incomplete visibility. Other propositions—AGI dates, discontinuous takeoff, digital consciousness, stable alignment, extinction probabilities—remain deeply contested.
The disciplined response is precision. Say “the study found” when describing its result, “the authors infer” when moving beyond it, “forecasters assign” when reporting belief, and “this scenario assumes” when exploring a pathway. Then link the source close to the claim.
That language may feel less dramatic than a countdown or a declaration. It is more useful. Humanity will make better decisions about advanced AI by knowing what the evidence supports, where it stops, and what would change the assessment.