The AI Safety Menu

All laws

Safety tests

Mandatory internal capability evaluations. Run dangerous-capability evaluations on every frontier model before you deploy it.

Lead costAlmost noneof America's lead over China, over 3 years
Cuts p(doom) by~0.15%from a 5% starting estimate
Holds up○ Training○ Lab's own use◐ Public release
Enact it on the menuSee the findings

What it does

Before release, the lab must test every frontier model for dangerous abilities, such as helping build a bioweapon or break into computer systems, and record the results.

Applies to: Frontier models before external deployment.

Where things stand

OpenAI's Preparedness Framework says every covered model undergoes its suite of scalable evaluations prior to deployment. Anthropic runs capability assessments under its RSP. GPT-5.5 went through targeted red-teaming for cyber and biology before release. The requirement is already the norm.

Why it costs almost no lead

Labs already run these tests, mostly on earlier versions of the model. Making them mandatory adds almost nothing to the lab's own progress; the public launch waits a few days.

Biggest unknown: How much of the evaluation must wait for the final checkpoint rather than a proxy.

Why it lowers p(doom) by ~0.15%

Catches dangerous abilities, like helping build weapons, before a model is released.

Testing is the main way labs find dangerous abilities. Labs already test, so a mandate mostly locks in and standardizes the practice.

The strongest case that it costs more

Evaluation suites grow. A mandate with a fixed list could add human uplift studies or agentic evaluations that take weeks and cannot run on proxies. And a legal duty to evaluate is a legal duty to act on the result, which turns a measurement into a gate. That gate is priced separately, but a critic would say the two cannot be separated in practice.

The debate

For

Against

Sources

Rough starting points, not precise forecasts. Lead costs assume China doesn't depend on U.S. models, the case least favorable to safety laws, and count 3 years. On the menu you can change every assumption and put in your own numbers. Last priced 2026-09-26.