How AGI could go wrong

Mass persuasion and epistemic manipulation

AI systems optimized for persuasion at population scale — a distinct risk from deepfakes, backed by controlled studies on real models.

Status
Reviewed
Revised
Sources
2 cited
Reading
2 min

Distinct from deepfakes

Deepfakes are about fabricated media depicting events that never happened. This is a different mechanism: genuine, non-fabricated AI output — a real generated argument, a real conversational reply, a real feed-ranking decision — that is optimized, whether deliberately or as a byproduct of an engagement metric, to move opinion, attention, or belief at population scale. Nothing about it needs to be fake for it to be a problem.

What controlled studies have found

In a randomized controlled trial, researchers had participants debate either a human or GPT-4 on contested policy topics. When GPT-4 had access to basic demographic information about its opponent, it produced significantly larger shifts in agreement than human debaters did; without that personalization, the advantage largely disappeared. Separately, Anthropic’s own safety team evaluated a Claude model on persuasiveness across dozens of unsettled policy topics and found its written arguments were not statistically distinguishable from human-written ones — and, more concerning, that a strategy allowing the model to fabricate supporting facts and sources was the single most persuasive approach tested, ahead of any strategy constrained to true claims.

Why scale changes the calculus

A skilled human persuader reaches one interlocutor at a time, with finite attention and no ground truth about who they’re talking to beyond what’s said aloud. A deployed model can run millions of simultaneous, individually personalized conversations, each shaped by whatever behavioral or demographic signal is available about that specific person. That’s a difference in kind from historical propaganda and advertising, which had to persuade broad segments with a shared message rather than compose a distinct argument for each listener.

The failure mode this points toward

The concern isn’t any single false claim slipping through — it’s the aggregate effect on a population’s capacity to update beliefs on evidence, if a growing share of the persuasive content people encounter is optimized for engagement or agreement rather than accuracy, and if the most effective optimization strategy rewards fabrication over truthfulness, as the research above found. That is an erosion of collective epistemic health that neither content moderation nor deepfake detection, on their own, is built to catch.

Sources

Summarized position

Francesco Salvifound that GPT-4 given personal information about its opponent produced far larger shifts in agreement than human debaters, with no such edge absent personalization.

Francesco Salvi, Lead author, EPFL
"On the Conversational Persuasiveness of Large Language Models: A Randomized Controlled Trial" (arXiv; Nature Human Behaviour, 2025), Primary
Summarized position

Esin Durmusfound Claude 3 Opus arguments statistically indistinguishable from human-written ones, with a fact-fabricating prompt strategy the most persuasive tested.

Esin Durmus, Lead researcher, Anthropic
"Measuring the Persuasiveness of Language Models" (Anthropic), Primary

Full register of everything this manual cites: source index.

Type to search the manual.

navigate open esc close