The AI Safety Menu

All laws

Power-grab safeguards

Safeguards against AI power grabs. Government AI from several providers, models that follow the law rather than one person, and audits for hidden backdoors.

Cuts lead by~2 daysof America's lead over China, over 3 years
Cuts p(doom) by~0.1%from a 5% starting estimate
Holds up○ Training◐ Lab's own use◐ Public release
Enact it on the menuSee the findings

What it does

Government and military AI must come from several providers, so no single company or person controls enough of it to seize power. Frontier models must follow the law rather than any one person's orders, and independent auditors get access to weights, training data and logs to check for hidden loyalties or backdoors. Reliable methods for detecting a well-hidden loyalty don't exist yet.

Applies to: Frontier AI systems, especially those used by governments and militaries.

Where things stand

Only voluntary lab policies exist. Anthropic's January 2026 constitution tells Claude to refuse to help concentrate power illegitimately, even if Anthropic asks, and lists helping seize absolute societal, military or economic control as a hard constraint; there is no outside audit for hidden loyalties.

Why it costs ~2 days

Audits of training data and model behaviour take staff time but don't hold up training. Under a day.

Biggest unknown: Whether audits can reliably detect a well-hidden loyalty at all; current interpretability tools may not be good enough, so a rule could give false assurance.

Why it lowers p(doom) by ~0.1%

Stops anyone secretly training AI to obey them instead of the law, a route to seizing power.

A small group controlling powerful AI is one path to a permanent power grab. Spreading government AI across providers and requiring models to follow the law make that harder today. Audits for hidden loyalties would add more, but detection methods are still unproven.

The strongest case that it costs more

Detecting a well-hidden backdoor may be impossible with current methods, so audits could give false assurance.

The debate

For

Against

Sources

Rough starting points, not precise forecasts. Lead costs assume China doesn't depend on U.S. models, the case least favorable to safety laws, and count 3 years. On the menu you can change every assumption and put in your own numbers. Last priced 2026-09-26.