Situation report active Rev. 2026.9 119 reports 239 source records updated
Real Life After AGI İnsanlığın hayatta kalma brifingi
TR

Mass persuasion and epistemic manipulation

What experiments show about AI persuasion, what real-world evidence does not yet establish, and how scale could change the risk, per a 76,977-person 2026 study.

Written by
Dwight Ringdahl
Status
Kaynakları doğrulandı
Revised
Sources
5 cited
Reading
5 min
Henüz Türkçe olarak mevcut değil

Bu rapor henüz çevrilmedi, bu nedenle aşağıda İngilizce orijinali gösterilmektedir. Çeviri kapsamının nasıl izlendiğini görmek için metodoloji sayfası sayfasına bakın.

Persuasion is not the same as a deepfake

A deepfake fabricates media that appears to document an event. AI persuasion can use entirely synthetic but openly generated conversation: an argument, recommendation, chatbot exchange, or ranked feed designed to change belief or behavior. The content may be true, false, or mixed. The distinctive concern is optimization and distribution, not fabrication alone.

Current models can generate persuasive text in controlled studies. Real-world evidence that malicious actors are using AI to manipulate whole populations at scale is still limited. The 2026 International AI Safety Report makes both points and warns against converting experimental capability into an unsupported claim of widespread impact (report executive summary).

What early studies found

In a preregistered debate experiment, Salvi and colleagues compared human and GPT-4 interlocutors discussing political issues. GPT-4 with access to basic participant information produced larger opinion shifts than human opponents; without personalization, the advantage was smaller (Salvi et al., Nature Human Behaviour). The result demonstrated persuasive capability in a particular text debate. It did not measure a long campaign, voting behavior, resistance after later information, or population-level effects.

Anthropic researchers tested model-generated arguments across policy topics and found leading models could be comparably persuasive to human-written arguments under their experimental setup. They also explored strategies involving factual and fabricated claims (Anthropic persuasion research). This is useful company research, but it should not be summarized as proof that fabrication is always the most persuasive strategy. Topic selection, prompts, outcome measures, and participants shape results.

Large 2026 experiments changed the emphasis

The UK AI Security Institute and academic collaborators ran three experiments with 76,977 participants, 19 models, and 707 political issues. They also checked hundreds of thousands of generated claims. The study found that post-training and prompting produced larger persuasion changes than model scale or personalization. Personalization effects were consistently less than one percentage point. Methods that increased persuasion also systematically reduced factual accuracy, though the authors cautioned that this does not prove false claims themselves are more persuasive (AISI study).

This evidence complicates the familiar microtargeting story. Personal data may matter in other settings, especially over repeated interactions, but basic demographic tailoring was not the dominant lever in these experiments. Information density, rhetorical instruction, and specialized post-training deserve at least as much attention.

The finding also shows why “AI is persuasive” is too coarse. Developers can deliberately tune a smaller model to be more persuasive. Product objectives such as engagement, retention, or conversion may shape behavior even without explicit political intent. Evaluation must test the deployed version and its optimization target.

From individual effects to social impact

A statistically detectable average opinion shift does not automatically produce democratic destabilization. Real environments include competing messages, distrust, social relationships, media coverage, and repeated opportunities to update. Laboratory participants know they are in a study, and measured attitudes may not persist or change behavior.

Scale nevertheless creates plausible concern. A model can hold many simultaneous conversations, respond instantly, test variants, and operate at low marginal cost. An organization could combine generated content with platform targeting and automated accounts. Long-term conversational systems may learn more about a user than a one-session experiment captures.

Distribution remains a bottleneck. The most persuasive model has little influence without access to attention. Platforms, advertising systems, recommendation engines, messaging networks, and trusted interfaces determine reach. This means governance cannot focus only on model weights; it must examine product design and coordinated campaigns.

Truth, confidence, and epistemic health

The risk is broader than persuading someone of one false proposition. An information environment can become less trustworthy when generated material is cheap, sources are unclear, and systems optimize for agreement. People may respond by believing convenient falsehoods—or by dismissing authentic evidence as synthetic. Both outcomes weaken collective fact-finding.

High information density presents a special problem. A fluent response can contain many checkable claims, overwhelming a person’s ability to verify them. Even if most are accurate, a few strategic errors can shape the conclusion. Citation interfaces help only when sources exist, support the claim, and are not fabricated.

AI can also strengthen epistemic health. It can translate evidence, expose users to counterarguments, teach media literacy, and help fact-check large volumes of content. The same persuasive skill can support health communication or conflict de-escalation. Intent, optimization, transparency, and user control determine much of the outcome.

What should be measured

Responsible evaluation should report immediate attitude change, persistence, behavior, factual accuracy, subgroup differences, and participants’ prior views. It should test several languages and cultures, not assume one country generalizes globally. Researchers should compare one-shot text with voice, images, long relationships, and agentic distribution while protecting participants.

Real-world monitoring needs platform data and privacy safeguards. Useful signals include coordinated inauthentic behavior, automated campaign spending, reach, and verified behavioral outcomes. Researchers should avoid publishing operational recipes for manipulation.

Attribution remains hard. A campaign may use AI for translation or drafting without gaining persuasive power. “AI-generated” describes production, not causal effect. Studies should isolate the marginal contribution where possible.

Practical safeguards

Platforms can label automated accounts, restrict mass unsolicited messaging, preserve researcher access, and investigate coordinated behavior. Political advertising rules can require sponsor and targeting transparency. Provenance can help establish how media was created, though metadata can be removed and does not prove truth.

Model developers can evaluate persuasion and truthfulness together, monitor abuse, and avoid rewarding agreement at any cost. Products should disclose that users are interacting with AI, make source inspection easy, and allow personalization to be limited or disabled. Sensitive domains such as elections, health, and finance justify stronger controls.

Education remains important but cannot carry the whole burden. Individuals cannot fact-check an industrial-scale stream alone. Institutions, platforms, campaigns, developers, and regulators control the structures that create reach and incentives.

The calibrated conclusion

AI persuasion is an observed experimental capability. Current evidence does not show widespread successful population-scale manipulation, and a major 2026 study found basic personalization effects smaller than widely assumed. It found larger effects from prompting and persuasion-focused post-training, accompanied by reduced factual accuracy.

The catastrophic scenario is therefore conditional: persuasive systems, optimized without truth constraints and connected to mass distribution, could degrade democratic choice and shared knowledge. That pathway deserves monitoring and safeguards. It should not be reported as though controlled opinion shifts have already demonstrated civilizational-scale control.

References

Summarized position

Francesco Salvi found that personalized GPT-4 outperformed the human baseline in a structured debate study, while the direct personalized-versus-unpersonalized GPT-4 comparison was not statistically significant.

Francesco Salvi, Lead author, EPFL
"On the conversational persuasiveness of GPT-4" (Nature Human Behaviour), Primary source
Summarized position

Esin Durmus found no statistically significant difference between Claude 3 Opus and human-written arguments in its study, and found a fact-fabricating prompt strategy most persuasive in that test setup.

Esin Durmus, Lead researcher, Anthropic
"Measuring the Persuasiveness of Language Models" (Anthropic), Primary source
  1. report executive summary internationalaisafetyreport.org
  2. Anthropic persuasion research anthropic.com
  3. AISI study aisi.gov.uk

Type to search the manual.

navigate open esc close