How to read this snapshot
This dashboard records public evidence available on September 13, 2026. It is not a live leaderboard and should not be silently updated. AI products, scaffolds, evaluation sets, and prices change too quickly for an undated “current” page to remain honest.
Capability is multidimensional. A system can solve an olympiad problem yet fail to read an analog clock, answer expert biology questions yet fail to complete a laboratory workflow, or pass a coding grader while producing a patch a maintainer would reject. Stanford’s 2026 AI Index calls this pattern jagged intelligence and documents large gaps between controlled and real environments (Stanford AI Index, technical performance).
The ratings below describe the evidence, not a universal model. “Strong” means leading systems perform impressively on relevant public evaluations. It does not mean safe, human-equivalent, or reliable enough for autonomous high-stakes use.
| Dimension | September 2026 evidence | Important limitation |
|---|---|---|
| Language and knowledge | Strong across broad question-answering and multimodal tests | Hallucinations, contamination, and uneven factual reliability remain |
| Mathematics and structured reasoning | Frontier systems reached top competition-level results on selected problems | Performance remains jagged and sensitive to tools and inference budget |
| Coding and computer use | Rapid gains on coding and structured desktop-agent benchmarks | Passing graders does not guarantee maintainable production work |
| Long-task autonomy | Measured software-task horizons increased rapidly | Mostly clean software/ML/cyber tasks; not duration of independent operation |
| Science assistance | Strong domain knowledge and improving protocol/troubleshooting performance | Knowledge scores do not establish end-to-end experimental competence |
| Cyber capability | Agents can find real vulnerabilities and complete selected expert tasks | End-to-end autonomous attacks are not publicly established |
| Persuasion | Measurable opinion shifts in controlled experiments | Population-scale malicious impact is not demonstrated |
| Robotics and embodiment | Strong progress in controlled environments | Household and open-world reliability remains much lower |
| Safety and reliability | Safeguards and evaluations are improving | Reporting is inconsistent and no method guarantees control |
Reasoning is impressive and uneven
The 2026 AI Index reports that a Google DeepMind system achieved a gold-medal-level score at the 2025 International Mathematical Olympiad under the competition time limit. The same chapter reports that the best tested model on ClockBench read analog clocks correctly only about half the time, versus roughly 90 percent for humans. These tests are not equally important, but their contrast defeats the idea of one scalar “intelligence level.”
Benchmark gains may reflect a better base model, more inference-time computation, tool use, sampling several solutions, or specialized scaffolding. A result belongs to the complete tested system under stated conditions. It should not be casually transferred to a cheaper product tier or a time-constrained deployment.
Many tests also saturate quickly. Stanford reports steep gains on Humanity’s Last Exam and near-saturation on SWE-bench Verified, while noting invalid questions and possible leaderboard adaptation. Saturation means an evaluation has less power to distinguish leading systems; it does not mean the underlying domain is solved.
Coding and computer use
Leading agents can repair selected repository issues, write substantial code, and operate graphical applications. Stanford reports OSWorld performance rising from about 12 percent to 66.3 percent, still failing roughly one in three structured tasks. Real computer use adds changing interfaces, unclear objectives, credentials, adversarial content, and consequences that a benchmark may omit.
Automated grading is another boundary. METR’s 2026 study of SWE-bench patches found that passing the benchmark’s tests did not mean maintainers would merge the changes, highlighting quality dimensions beyond functional test cases (METR maintainer review study). Production coding requires architecture, communication, security, maintenance, and responsibility for failures.
This evidence supports substantial assistance and automation of selected work. It does not support “software engineering is solved.”
Long-task autonomy
METR’s May 2026 dashboard estimates the human-expert task duration at which an agent has a 50- or 80-percent predicted success rate. The tasks primarily concern software engineering, machine learning, and cybersecurity. METR states that measurements above sixteen hours are unreliable with the current suite (METR task horizons).
The metric is often misreported. A four-hour horizon does not mean the agent runs independently for four hours; it means a human expert typically needs about four hours for tasks of the measured difficulty. Successful agents may finish much faster. The suite is self-contained and objectively gradable, unlike much organizational work.
METR further cautions that domains differ by orders of magnitude, error bars can be wide, and a 50-percent horizon is inadequate where 98-percent reliability is required (METR limitations). The trend is meaningful evidence of rapid progress in selected digital tasks—not a clock counting down to job automation or AGI.
Science, biology, and cyber
The UK AI Security Institute reports that systems tested through October 2025 exceeded its PhD-expert baselines on selected chemistry and biology questions and improved at protocol generation and laboratory troubleshooting. It also reports persistent difficulty with end-to-end plasmid design (AISI Frontier AI Trends Report). These are government evaluation results, but expert baselines and task construction still determine meaning.
Cyber capability has advanced from question answering toward agents that discover vulnerabilities and solve multi-step challenges. AISI reported faster growth in its cyber-task measures in February 2026 (AISI cyber trends). The International AI Safety Report distinguishes that evidence from deployment: attackers use AI, yet fully autonomous end-to-end attacks have not been publicly reported and whether offense or defense gains more remains uncertain (International AI Safety Report 2026).
Persuasion and social capability
Conversational AI can change attitudes in experiments. A 2026 AISI-led study involving 76,977 participants found larger persuasion effects from prompting and persuasion-focused post-training than from scale or basic personalization; personalization effects were below one percentage point. Increased persuasion was associated with reduced factual accuracy (AISI persuasion study).
That finding establishes measured capability under experimental conditions. It does not prove durable behavior change, election effects, or population control. Reach, platform distribution, competing messages, user trust, and persistence remain separate variables.
Embodiment remains a major gap
Robotics shows the clearest difference between controlled and open environments. Stanford reports high success in a simulation benchmark but only 12 percent on tested real household tasks. Autonomous vehicles operate at meaningful scale in constrained service areas, with operational design domains and off-site human support that matter to interpretation.
A language model’s broad digital interface should not be mistaken for physical-world generality. Sensing, manipulation, safety around people, recovery from novel errors, hardware maintenance, and cost all constrain deployment. This manual’s humanoid-robots chapter tracks that gap between benchmark demonstrations and real-world robot performance in more detail, company by company.
Reliability is part of capability
An agent that succeeds 66 percent of the time may be useful with review and unacceptable without it. High-stakes systems require calibrated uncertainty, failure detection, appeal, audit logs, rollback, security, and human fallback. Average benchmark scores hide rare but consequential errors.
The 2026 AI Index finds that developers report mainstream capability benchmarks far more consistently than responsible-AI evaluations (AI Index responsible-AI chapter). Missing public results do not prove absent internal testing, but they limit comparison and public assurance.
What this dashboard does not establish
It does not establish consciousness, intentions, economic value across all jobs, or an accepted AGI threshold. It does not show that a benchmark result transfers to a different language, population, toolset, cost limit, or adversarial environment. It does not resolve whether current architectures can become robustly general.
As of this snapshot, the strongest conclusion is that frontier systems show extraordinary breadth and rapid improvement alongside persistent brittleness, domain gaps, and incomplete safety evidence. That is more consequential—and more accurate—than either “only autocomplete” or “AGI has arrived.”