Situation report active Rev. 2026.9 119 reports 239 source records updated
Real Life After AGI 人類存続のためのブリーフィング
JA

Whistleblower protections inside AI labs

How whistleblower rights, reporting channels and anti-retaliation rules can help AI-lab employees expose safety failures before they become public harms.

Written by
Dwight Ringdahl
Status
出典確認済み
Revised
Sources
4 cited
Reading
5 min
日本語ではまだ提供されていません

このレポートはまだ翻訳されていないため、以下に英語の原文を表示しています。翻訳範囲がどのように追跡されているかについては 方法論のページ をご覧ください。

Insiders see evidence before the public

Employees and contractors may encounter dangerous evaluation results, security failures, misleading public claims, unlawful data practices, or pressure to deploy before concerns are resolved. Regulators and customers often see only curated reports. Protected internal and external reporting can shorten the time between discovery and correction.

This is not a hypothetical concern. In 2024, current and former employees of OpenAI and Google DeepMind publicly called on AI companies not to retaliate against staff who raise safety concerns after internal channels fail — a demand endorsed by three of the field’s most cited safety researchers, Yoshua Bengio, Geoffrey Hinton, and Stuart Russell (current and former OpenAI and Google DeepMind employees, 2024). That the demand had to be made publicly, by name, is itself evidence that existing internal channels were not seen as sufficient.

Whistleblowing is not a substitute for audits or competent management. Employees can be mistaken, lack context, or act from mixed motives. A good system protects good-faith reporting and independently investigates evidence rather than assuming either the company or the complainant is right.

There is no single U.S. federal “AI whistleblower law” covering every safety concern. Existing protections depend on employer, industry, contract, jurisdiction, and what is reported. Laws may protect disclosures about securities fraud, government contracts, discrimination, workplace safety, or specific violations, while a concern about a model’s catastrophic capability may not fit cleanly.

Nondisclosure agreements cannot lawfully erase every protected right, but employees may not know the boundary and may face expensive disputes. Trade-secret, privacy, and national-security rules can restrict public disclosure even when a protected channel to a regulator should exist. This page is general information, not legal advice.

California created AI-specific protection

California’s Transparency in Frontier Artificial Intelligence Act, SB 53, took effect in 2026. Within its definitions, covered employees can report violations or a specific and substantial danger to public health or safety resulting from catastrophic risk. The California Attorney General operates a reporting channel and explains coverage (California Department of Justice, 2026).

SB 53 also requires covered large frontier developers to publish and follow frontier AI frameworks and report specified critical safety incidents. It is a meaningful state mechanism, but not universal protection: definitions of developer, employee, model, and catastrophic risk determine coverage. Readers should consult the enacted text and qualified counsel for a real case.

New York’s RAISE Act, signed in December 2025, likewise created frontier-model safety and reporting requirements within New York jurisdiction (New York Governor). State regimes may converge, conflict, or face litigation; current status must be checked.

The European Union added a reporting channel

The EU AI Act enforcement framework includes a secure whistleblower tool for people professionally connected to AI-system or general-purpose-model providers to report suspected violations. The AI Office and national authorities divide enforcement responsibility (European Commission, August 2026).

That tool concerns violations of applicable EU law, not every ethical disagreement or speculative risk. EU whistleblower protections and national implementation also affect remedies. A channel is only effective if users understand jurisdiction, confidentiality, evidence handling, and protection from retaliation.

What effective internal reporting requires

An employee should have several routes: a manager, an independent safety or compliance function, a board committee, and an authorized external regulator. Reporting through the same executive chain that approved a disputed release creates a conflict. Safety leaders need access to the board and protection from commercial retaliation.

Policies should define prohibited retaliation broadly: firing, demotion, lost equity, undesirable assignments, threats, blacklisting, immigration pressure, and strategic litigation. Contractors, temporary workers, red-team vendors, and researchers often need coverage because they may see the same evidence as employees.

Acknowledgment deadlines, preservation duties, independent triage, and written outcomes make a channel auditable. Anonymous reporting helps, but anonymity can fail through technical metadata or the small number of people with access to an event. Organizations should minimize access and explain limits honestly.

Equity and compensation can silence dissent

Frontier-lab compensation may include valuable unvested equity. Losing it after resignation can make speaking up financially ruinous even without a direct threat. Separation agreements can add nondisparagement, confidentiality, or non-disclosure pressure.

A fair system should preserve already-earned compensation, provide time and independent advice before signing releases, state protected-reporting rights clearly, and prohibit contractual terms that condition economic value on silence about lawful safety concerns. These protections reduce coercion without authorizing disclosure of unrelated secrets.

Evidence handling and responsible escalation

Potential whistleblowers should use authorized channels and obtain legal advice before copying confidential data. Mass removal of user records, model weights, personal information, or classified material can create new harms and legal exposure. The goal is to preserve necessary evidence while minimizing unrelated disclosure.

Organizations should maintain tamper-evident logs, evaluation records, deployment decisions, dissenting analyses, and incident timelines. Retention prevents a concern from becoming one person’s word against another. Regulators need secure facilities and technical staff capable of examining evidence that cannot be made public.

Public disclosure may sometimes be justified, but it is not the first safe option in every case. A mature system offers credible independent channels before employees feel forced to choose between silence and releasing sensitive material to the world.

Avoiding malicious or low-quality reports

Protection for good-faith reporting does not require accepting fabricated evidence or harassment. Triage can evaluate specificity, corroboration, access, and urgency. Knowingly false reports can remain sanctionable, but disagreement or an unconfirmed concern should not be treated as bad faith.

Feedback loops matter. If employees see reports disappear without explanation, they will stop using the channel. Aggregate transparency—number of concerns, categories, investigation time, substantiation, and corrective action—can demonstrate function while protecting identities.

Why voluntary promises are insufficient

Company policies can exceed legal requirements and adapt quickly. They can also change, omit contractors, or be enforced by people with conflicting incentives. Legal rights provide an external floor, remedies, and subpoena power. The strongest architecture combines internal culture, board responsibility, regulator channels, and enforceable anti-retaliation rules.

Laboratories are interested parties in describing their own safety culture. Departing employees are also interested parties with partial views. Independent investigations should evaluate documents and behavior. A pattern of departures is a warning signal, not proof of a particular allegation.

A model policy agenda

Governments could protect good-faith disclosures about defined serious AI risks even before a law is violated; cover employees and contractors; preserve access to regulators and legislatures; provide confidentiality, interim relief, reinstatement, damages, and fee recovery; and punish knowing retaliation. Rules should coordinate with security and privacy law.

Regulators should publish clear intake criteria, protect sensitive evidence, employ technical experts, and report anonymized outcomes. Boards should certify that significant safety concerns reached independent directors before high-risk deployment.

The bottom line

Whistleblower protection is relatively low-cost compared with training or auditing frontier models, but it is not automatic. California, New York, and the EU now provide concrete AI-related mechanisms within their jurisdictions; coverage elsewhere remains fragmented. The measure of success is not a policy page. It is whether people can raise specific evidence, reach an independent investigator, preserve their livelihood and lawful rights, and trigger correction before harm scales.

References

Summarized position

Current and former OpenAI and Google DeepMind employees calls on AI companies not to retaliate against employees who publicly raise safety concerns after other channels fail.

Current and former OpenAI and Google DeepMind employees, "A Right to Warn about Advanced Artificial Intelligence," endorsed by Yoshua Bengio, Geoffrey Hinton, and Stuart Russell
righttowarn.ai, Signed statement
  1. California Department of Justice, 2026 oag.ca.gov
  2. New York Governor governor.ny.gov
  3. European Commission, August 2026 digital-strategy.ec.europa.eu

Type to search the manual.

navigate open esc close