Power-grab safeguards
Safeguards against AI power grabs. Government AI from several providers, models that follow the law rather than one person, and audits for hidden backdoors.
What it does
Government and military AI must come from several providers, so no single company or person controls enough of it to seize power. Frontier models must follow the law rather than any one person's orders, and independent auditors get access to weights, training data and logs to check for hidden loyalties or backdoors. Reliable methods for detecting a well-hidden loyalty don't exist yet.
Applies to: Frontier AI systems, especially those used by governments and militaries.
- Government and military AI must be bought from several independent providers.
- Frontier models must follow the law rather than any one person's orders.
- Independent auditors get access to weights, training data and logs to check for hidden loyalties and backdoors.
Where things stand
Only voluntary lab policies exist. Anthropic's January 2026 constitution tells Claude to refuse to help concentrate power illegitimately, even if Anthropic asks, and lists helping seize absolute societal, military or economic control as a hard constraint; there is no outside audit for hidden loyalties.
Why it costs ~2 days
Audits of training data and model behaviour take staff time but don't hold up training. Under a day.
Biggest unknown: Whether audits can reliably detect a well-hidden loyalty at all; current interpretability tools may not be good enough, so a rule could give false assurance.
Why it lowers p(doom) by ~0.1%
Stops anyone secretly training AI to obey them instead of the law, a route to seizing power.
A small group controlling powerful AI is one path to a permanent power grab. Spreading government AI across providers and requiring models to follow the law make that harder today. Audits for hidden loyalties would add more, but detection methods are still unproven.
The strongest case that it costs more
Detecting a well-hidden backdoor may be impossible with current methods, so audits could give false assurance.
The debate
For
- Tom Davidson, Lukas Finnveden, Rose Hadshar (Forethought) 2025 report urging rules against secret loyalties, alignment audits with full access to internals and training data, and multiple independent providers for military AI
- Joe Kwon, Tom Davidson, Owain Evans, Markus Anderljung, Ryan Greenblatt, Daniel Kokotajlo and others 2026 ICML workshop paper: secret loyalties are a serious but addressable threat; proof-of-concept secret loyalties already evade black-box audits in open-weight models
- Foundation for American Innovation (Govind Pimpale, Blaine Dillingham) 2026 proposal for an independent evaluator (e.g., in GAO) with power to red-team AI used across government for hidden loyalties
- Anthropic January 2026 constitution: Claude should refuse to help concentrate power illegitimately, even at Anthropic's request
Against
- NetChoice 2026 testimony against Illinois SB 315: mandatory third-party frontier audits are an impossible burden with no auditing standards and risk trade secrets
- President Donald Trump August 2026 remarks that Congress wants to regulate AI 'out of business', read as opposition to the FRONTIER Act's mandatory audits
Sources
- AI-Enabled Coups: How a Small Group Could Use AI to Seize Power (Forethought, April 2025): Core paper; names singular loyalty, secret loyalties and exclusive access as the three risks
- AIs with Secret Loyalties are a Serious but Addressable Threat (ICML 2026): Technical evidence and research agenda
- Secret Loyalties in Government AI (FAI, July 2026): Policy proposal for Congress
- Tom Davidson on AI-enabled coups (80,000 Hours podcast): Accessible explainer, covers reasons for skepticism
Rough starting points, not precise forecasts. Lead costs assume China doesn't depend on U.S. models, the case least favorable to safety laws, and count 3 years. On the menu you can change every assumption and put in your own numbers. Last priced 2026-09-26.