Situation report active Rev. 2026.4 119 reports 237 source records updated
Real Life After AGI The human survival briefing

Frontier AI labs' safety frameworks, compared

Six labs' safety frameworks compared: thresholds, safeguards, and what happens when one is crossed, checked against independent audits of actual practice.

Written by
Dwight Ringdahl
Status
Reviewed
Revised
Sources
13 cited
Reading
12 min

What a safety framework actually is

A frontier AI lab’s safety framework is the lab’s own published policy defining capability thresholds — specific things a model might be able to do — and the safeguards or organizational commitments meant to trigger once a model reaches one of them. Anthropic calls its version a Responsible Scaling Policy; OpenAI, a Preparedness Framework; Google DeepMind, a Frontier Safety Framework; Meta, an Advanced AI Scaling Framework; xAI, a Frontier Artificial Intelligence Framework; Microsoft, a Frontier Governance Framework. The names differ; structurally, all six try to answer the same question — how does a company know when a model is too dangerous to deploy as-is, and what is it actually committed to doing about it?

That question matters because no external regulator currently sets binding, model-specific capability thresholds for any of these six companies (see AI law and enforcement in 2026). For now, these documents are the industry’s real working answer. They are also not the same document copied six times: which risks each lab tracks, how a threshold is defined, and — most consequentially — what happens when one is crossed differ in ways that matter.

One thing needs to be said before the comparison starts, because it shapes how to read everything below: every claim in the table is the lab’s own stated policy. None of it is independent verification that the lab actually does what it says. A separate section after the table, “Claims versus practice,” covers the graded external audits, academic critiques, and documented gaps between these six companies’ stated commitments and their actual conduct — and that section, not the table, is the more important one.

Scope and sourcing note

Every framework detail below was checked against the lab's own current published document, plus (for the "claims versus practice" section) the independent reports and papers named, as of September 2026. Exact version numbers and effective dates are given throughout so a returning reader can tell whether something has changed. These policies are revised often — Anthropic alone has published nine numbered versions since September 2023 — so treat every row as a dated snapshot, not a settled description.

The frameworks, compared

Lab (framework) Current version / date Risk domains tracked Threshold / tier structure What happens when a threshold is crossed
Anthropic — Responsible Scaling Policy v3.4, effective Jul 8, 2026 Non-novel CBRN uplift; novel CBRN uplift; high-stakes sabotage; automated AI R&D acceleration Capability/usage-threshold table; fixed ASL-2/ASL-3/ASL-4 tiers dropped in the v3.0 rewrite (legacy ASL labels kept only for present-day controls already in place) Public, non-binding Frontier Safety Roadmap, graded against past performance, plus Risk Reports every 3–6 months with external review — no unconditional pause
OpenAI — Preparedness Framework v2, Apr 15, 2025 Biological & Chemical, Cybersecurity, AI Self-Improvement (tracked); Long-range Autonomy, Sandbagging, Autonomous Replication and Adaptation, Undermining Safeguards, Nuclear & Radiological (research categories, not yet full criteria); Persuasion explicitly excluded Two tiers: High and Critical High: safeguards required before deployment. Critical: safeguards required during development; can halt development entirely. Internal Safety Advisory Group recommends; CEO/designee decides; board Safety and Security Committee oversees
Google DeepMind — Frontier Safety Framework v3.1, Apr 17, 2026 CBRN; Cyber; Harmful Manipulation (flagged “exploratory and subject to further research”); ML R&D and Misalignment (merged into one domain in v3.1) Critical Capability Levels (CCLs) plus lower-bar Tracked Capability Levels (TCLs, new in v3.0/3.1); Security Levels SL1–SL4 used as vocabulary, not a compliance standard An “alert threshold” triggers a formal capability assessment before a CCL is reached; reaching a CCL requires a supplemental safety case and governance sign-off before external deployment (and, for ML R&D CCLs specifically, before high-risk internal deployment too)
Meta — Advanced AI Scaling Framework v2, Feb 3, 2025 (renamed from “Frontier AI Framework” v1.1; superseded roughly Apr 2026) Chemical & Biological; Cybersecurity; Loss of Control (added in v2) “Level of uplift toward a threat scenario,” not a fixed tier scale If critical risk can’t be mitigated: will not deploy the model externally or continue its development
xAI — Frontier Artificial Intelligence Framework (FAIF) Effective Jun 30, 2026 (superseded the earlier Risk Management Framework, last updated Aug 20, 2025) CBRN; Offensive Cybersecurity; Loss of Control; Harmful Manipulation Full systemic risk assessment at least annually; no fixed tier scale Conditional language throughout (“we may,” “if we determine it is warranted”); no explicit halt-development trigger comparable to peers
Microsoft — Frontier Governance Framework Feb 2026 update (v1 was Feb 2025) CBRN weapons; Offensive cyberoperations; Advanced autonomy; Loss of control and Harmful manipulation (added Feb 2026) Two-stage: automated “leading indicator” benchmark screening at least every 6 months, escalating to deeper assessment rated Low/Medium/High/Critical “If… we identify a risk we cannot sufficiently mitigate, we will pause development and deployment until the point at which mitigation practices evolve to meet the risk” — the most direct halt-commitment among all six

The most consequential change: Anthropic’s move off an unconditional pause

Anthropic’s RSP has the longest public history of the six, and its most recent structural change is the single most-discussed development among frontier safety frameworks in 2026. Version 1.0, published September 19, 2023, introduced the AI Safety Level (ASL) system — ASL-2, ASL-3, ASL-4 — each defined by a specific, fixed list of required controls, alongside a hard commitment: Anthropic would not train past a threshold without safeguards already proven adequate in advance. Versions 2.0 (October 2024), 2.1 (March 2025), and 2.2 (May 2025) extended that structure and added the CBRN threshold.

Version 3.0, effective February 24, 2026, rewrote the policy. Anthropic replaced the fixed ASL tiers with a capability/usage-threshold table, and explained the reasoning directly in an appendix — the company’s own words on why it walked away from specific, itemized control lists appear below as a registered citation. In place of the old hard pause, v3.0 introduced two new instruments: a public, non-binding Frontier Safety Roadmap graded openly against Anthropic’s own past performance, and Risk Reports published every three to six months and subject to review by at least one vetted, external, conflict-free reviewer. Anthropic’s document is explicit that this is a deliberate shift from a unilateral, absolute-risk posture to one that is conditional and competitor-contingent: a named appendix lays out scenarios such as “Anthropic in the lead” and “Competitors have strong safety measures,” under which the company’s own commitments change. Versions 3.1 (April 2), 3.2 (April 29), 3.3 (May 26), and 3.4 (July 8, 2026, current) have each revised details since — most recently narrowing the automated-R&D threshold and widening internal Risk Report distribution — without restoring the earlier hard pause.

Governance around the policy itself has stayed constant across these versions: a named Responsible Scaling Officer, a noncompliance-reporting channel, and an annual third-party procedural compliance review that Anthropic itself describes as checking “procedural compliance, not substantive outcomes” — a distinction worth noticing, since it means the audit confirms the process was followed, not that the resulting decisions were correct.

The change registered immediately as a credibility event outside the company. Chris Painter, METR’s Director of Policy, reviewed an early draft with Anthropic’s permission and, while welcoming the added transparency of the Risk Reports, gave reporters the blunt assessment quoted below.

Where the other five frameworks converge and diverge

Microsoft’s Frontier Governance Framework is the most direct contrast to Anthropic’s new conditional posture: its pause commitment, quoted in full in the table above and below, is unconditional on its face, tied only to Microsoft’s own assessment that a risk can’t be sufficiently mitigated — no competitor-contingent scenario carve-out comparable to Anthropic’s Appendix A.

Meta’s Advanced AI Scaling Framework makes a similarly unconditional commitment for its narrower case: if a model reaches Meta’s “critical risk threshold” and that risk can’t be mitigated, Meta states it will not deploy the model externally or continue its development. Meta uses a “level of uplift toward a threat scenario” rather than a fixed tier scale, but the underlying domains — chemical/biological risk, cybersecurity, and Loss of Control (added in Meta’s v2) — turn out to be more shared across labs than each framework’s separate publication implies. OpenAI’s own Preparedness Framework v2 says outright, in a footnote, that its Capability Reports and Safeguards Reports “parallel Anthropic’s updated RSP” and that its threshold criteria “were informed in part by Meta’s recent Frontier AI Framework” — a rare acknowledgment of cross-lab convergence in a set of documents that otherwise read as though each were written in isolation.

OpenAI’s own two-tier structure puts the harder gate — a possible halt to development itself — only at the Critical tier, and even then routes the decision through an internal Safety Advisory Group whose recommendation OpenAI’s CEO or a designee can override, with the board’s own Safety and Security Committee (co-led by that same CEO) providing oversight of the override. The framework was applied in practice in July 2026: OpenAI’s GPT-5.6 System Card designated its smaller, faster Sol, Terra, and Luna models as High capability in both Cybersecurity and Biological and Chemical risk — the first time smaller models in a family received that designation — triggering new activation classifiers, real-time output blocking, and roughly 700,000 GPU-hours of automated red-teaming (OpenAI, GPT-5.6 deployment safety page). That is OpenAI’s own account of its own compliance, not an independent audit.

Google DeepMind’s Frontier Safety Framework works differently again: rather than a binary tier crossing, an “alert threshold” is meant to trigger a formal capability assessment before a Critical Capability Level is actually reached, and reaching one requires a supplemental safety case and governance sign-off before external deployment — extended, for the ML R&D domain specifically, to high-risk internal deployment as well (DeepMind, Frontier Safety Framework v3.1). The clearest example of a lab naming the limits of its own framework in public: DeepMind’s recommended Security Level 4 for models that could fully automate an entire Google AI research team’s work comes with the explicit note that this “must be taken on by the frontier AI field as a whole” — DeepMind stating, in its own published document, that no single company’s controls are sufficient at that capability level.

xAI’s FAIF is comparatively thin on this dimension. Across its four domains, conditional language is nearly constant — “we may,” “if we determine it is warranted” — and the document contains no deployment-halting commitment comparable to OpenAI’s Critical tier or Meta’s “will not deploy” language (xAI, Frontier Artificial Intelligence Framework). It replaced an earlier Risk Management Framework, itself already published months behind xAI’s own stated schedule — a pattern with a documented history, covered next.

For a broader index of published frontier-lab policies beyond these six, METR maintains a comparison (METR, “Common Elements of Frontier AI Safety Policies”; index) covering twelve labs. It functions mainly as a structural index rather than a graded scorecard — for a graded comparison, see the next section.

Claims versus practice

Everything above describes what six companies say they will do. Whether they have actually done it is a different, and more important, question, and the evidence available is considerably less flattering than the frameworks themselves.

The single most useful independent check is the SaferAI Frontier Risk Management Tracker, a genuinely independent rating distinct from any lab’s self-reported framework. It grades twelve companies — not only the six compared above — on four dimensions: Risk Identification, Risk Analysis & Evaluation, Risk Treatment, and Risk Governance. As of its July 2026 update, overall scores out of 100% were: Anthropic 35% (highest overall, and strongest on Risk Governance specifically at 50%), OpenAI 34%, Microsoft 33%, Meta 33%, G42 24%, Google DeepMind 20%, xAI 18%, Amazon 18%, NVIDIA 16%, Magic 11%, Naver 10%, and Cohere 8% (lowest) (SaferAI Frontier Risk Management Tracker). SaferAI’s own benchmark for what current best practice would produce, if every company adopted the single strongest existing practice found anywhere across the field, is quoted below — and it is nowhere close to 100%. No lab reviewed, including the leader, is close to what SaferAI considers adequate.

OpenAI’s Preparedness Framework has been checked directly by outside academics rather than taken at face value. Coggins, Saeri, Daniell, Ruster, Liu, and Davis apply an “affordances” framework — distinguishing what a policy document demands, requests, encourages, and merely allows — and reach three findings. First, the Framework requests evaluation of only a minority of the risk subdomains in the MIT AI Risk Repository’s 24-subdomain taxonomy, and demands evaluation of none. Second, it encourages deployment of systems with “Medium” capability for severe harm — OpenAI’s own definition covers the death or grave injury of thousands of people, or hundreds of billions of dollars in economic damage — citing OpenAI’s own deployment of its o1 model despite Medium-rated biological/chemical and persuasion capability. Third, it allows OpenAI’s CEO to override the Safety Advisory Group’s safeguard recommendations, while that same CEO co-leads the board committee meant to oversee that discretion — a structural conflict-of-interest finding, not a hypothetical one.

Google DeepMind’s own safety case has been externally audited — the clearest example, among the six labs, of a specific safety-relevant claim being checked by named, credentialed outside reviewers rather than the lab’s internal process. DeepMind’s internal argument that its models lack the ability to “scheme” was reviewed by Barrett, Campos Zabala, Fillingham, Siddique, Walpole, Bloomfield, and Papadatos using a formal assurance methodology. Their assessment credited DeepMind with genuine methodological engagement and transparent reasoning, but found what is quoted below — a finding about the distance between engineering practice and the evidence actually presented, not an accusation of bad faith.

xAI supplies the clearest case of a framework’s claims being directly contradicted by events, and it happened twice. First, on process: xAI’s February 2025 draft framework promised a finalized version “within three months” — a deadline of roughly May 10, 2025. The deadline passed with no acknowledgment from xAI, prompting the watchdog group The Midas Project to note the pattern quoted below. The eventual FAIF predecessor wasn’t published until August 20, 2025 — over three months late. Second, on substance: that same day, xAI accidentally published hundreds of thousands of user conversations, and Elon Musk separately acknowledged that an engineer had downloaded the company’s entire code repository — directly undercutting the same document’s own written security claims. AI Lab Watch also documented that third-party evaluator credit for the UK AI Security Institute was quietly removed from the Grok 4 model card one day after publication, without explanation (AI Lab Watch, “xAI’s new safety framework is dreadful”).

Not every entry in this section is a documented failure. OpenAI’s GPT-5.6 High-capability designation, described above, is best read as a framework being applied largely as written — but it remains OpenAI’s own self-report; no independent verification of the red-teaming or classifier-effectiveness claims behind it was found as of this writing.

How to use this page

Every framework named here gets revised on its own schedule, usually without much public notice beyond a changelog entry. The version numbers and effective dates above are the fastest way to check whether a given row is still current: Anthropic’s RSP is at v3.4 (July 8, 2026), OpenAI’s Preparedness Framework has been at v2 since April 15, 2025 with no v2.1 as of this writing, DeepMind’s Frontier Safety Framework is at v3.1 (April 17, 2026), Meta’s Advanced AI Scaling Framework is at v2 (February 3, 2025, since superseded), xAI’s FAIF took effect June 30, 2026, and Microsoft’s Frontier Governance Framework was last updated February 2026. If any of those numbers have moved by the time you’re reading this, treat the table as stale on that row specifically.

For how the capability tests and safety-case arguments behind these frameworks actually work — and their own limits — see frontier evaluations, safety cases, and incident reporting. For a chronological record of the real-world incidents, disclosures, and near misses that these frameworks are meant to prevent or catch, see the AI incidents timeline.

References

Direct quotation
“Earlier editions of our RSP defined “AI Safety Levels” with specific lists of required controls... when defining the risk mitigations needed for future levels of AI capability, we have found that providing a specific list of controls is overly rigid, and we instead prefer to focus on what sort of argument an AI developer should make”
Anthropic, Responsible Scaling Policy, Appendix B ("Notes on ASLs")
anthropic.com, Primary
Direct quotation
“This is more evidence that society is not prepared for the potential catastrophic risks posed by AI”
Chris Painter, Director of Policy, METR
Winbuzzer, Interview
Direct quotation
“If, during the implementation of this framework, we identify a risk we cannot sufficiently mitigate, we will pause development and deployment until the point at which mitigation practices evolve to meet the risk”
Microsoft, Frontier Governance Framework (February 2026 update)
microsoft.com, Primary
Direct quotation
“if a company were to apply all the best practices currently found across the other companies, they would achieve a score of 59%”
SaferAI, Frontier Risk Management Tracker
tracker.safer-ai.org, Report
Direct quotation
“OpenAI's Framework requests research & evaluation of a small minority of AI risks and demands none”
Sam Coggins, Alexander K. Saeri, Katherine A. Daniell, Lorenn P. Ruster, Jessie Liu, and Jenny L. Davis, Authors, "The 2025 OpenAI Preparedness Framework does not guarantee any AI risk mitigation practices: a proof-of-concept for affordance analyses of AI safety policies"
arXiv:2509.24394, Report
Direct quotation
“we identified gaps between what would be expected in established safety-engineering practice and what the safety case delivers”
Stephen Barrett, Francisco Javier Campos Zabala, Sean P. Fillingham, Umair Siddique, James Walpole, Robin Bloomfield, and Henry Papadatos, Authors, "Lessons from External Review of DeepMind’s Scheming Inability Safety Case"
arXiv:2604.21964, Report
Direct quotation
“xAI misses a second self-imposed deadline to implement a frontier safety policy”
The Midas Project, AI-safety watchdog group, quoted by TechCrunch
TechCrunch, Secondary
  1. OpenAI, GPT-5.6 deployment safety page deploymentsafety.openai.com
  2. DeepMind, Frontier Safety Framework v3.1 storage.googleapis.com
  3. xAI, Frontier Artificial Intelligence Framework media.x.ai
  4. METR, "Common Elements of Frontier AI Safety Policies" metr.org
  5. index metr.org
  6. AI Lab Watch, "xAI's new safety framework is dreadful" ailabwatch.substack.com

The source index also tracks the manual's recurring core sources and expert positions.

Type to search the manual.

navigate open esc close