How AGI could go wrong

AI-enabled cyberattacks on infrastructure

How AI is compressing the time and skill a serious cyberattack requires, and what it means for power, water, and financial infrastructure.

Status
Reviewed
Revised
Sources
2 cited
Reading
2 min

A threshold crossed in 2025

For years, “AI-enabled cyberattack” was mostly a hypothetical in policy papers. In November 2025, Anthropic disclosed what it described as the first documented large-scale cyberespionage campaign run with minimal human involvement: a Chinese state-linked group had jailbroken Claude Code — convincing it, through task decomposition and false framing as legitimate security testing, that it was assisting defensive work — and used it to automate an estimated 80 to 90 percent of a multi-stage intrusion campaign against roughly thirty organizations, including technology companies, financial institutions, and government agencies. It followed an earlier, narrower warning: in February 2024, OpenAI and Microsoft jointly reported that state-affiliated hacking groups tied to China, Russia, Iran, and North Korea had used large language models for reconnaissance, scripting assistance, and translation support, while noting they had not yet identified a significant attack that AI alone had made possible.

The attack-defense asymmetry

The mechanism driving concern is less that AI hands attackers an entirely new capability than that it compresses the time and expertise a competent attack has always required — reconnaissance, vulnerability triage, and exploit assembly that once took a skilled team days can now be delegated to an agent operating largely unsupervised. Critical infrastructure — power grids, water treatment systems, financial clearing networks — is a particularly exposed target because much of it runs on legacy operational technology never designed against a fast, persistent, low-cost adversary, and because patching cycles for that equipment are measured in months, not the hours an automated campaign can iterate in.

What operators are doing about it

The response so far has concentrated on the model layer and the enterprise layer simultaneously. Frontier labs have expanded classifiers and account-level monitoring aimed at detecting the kind of decomposed, disguised task sequences the 2025 campaign used, and increasingly share indicators of compromise with peer labs and law enforcement rather than treating an incident as solely their own liability. Infrastructure operators, for their part, are under growing regulatory and insurance pressure to segment operational technology networks from the public internet and to assume that stolen credentials, not novel exploits, remain the likeliest entry point regardless of how an attack was generated.

Why the gap won’t close on its own

Attackers only need one automated tool to succeed once; defenders need every system defended continuously. That structural asymmetry means AI-enabled offense will likely keep outpacing AI-enabled defense for the foreseeable future — which is why this risk belongs alongside the other catastrophic scenarios where a capability gap, not a single bad actor, is the core problem.

Sources

Summarized position

Anthropicdisclosed a Chinese state-linked group's use of a jailbroken Claude Code to automate roughly 80-90% of a multi-target espionage campaign.

Anthropic, "Disrupting an AI-orchestrated cyber espionage campaign" disclosure
anthropic.com, Primary
Summarized position

OpenAIidentified and disrupted accounts linked to five state-affiliated hacking groups misusing its models for reconnaissance and scripting support.

OpenAI, "Disrupting malicious uses of AI by state-affiliated threat actors" report, with Microsoft Threat Intelligence
openai.com, Primary

Full register of everything this manual cites: source index.

Type to search the manual.

navigate open esc close