The AI Safety Menu

All laws

Red team if risky

Triggered red-team and remediation. If an evaluation crosses a danger threshold, do elevated testing and fix it before anyone can use it.

Lead costAlmost noneof America's lead over China, over 3 years
Cuts p(doom) by~0.15%from a 5% starting estimate
Holds up○ Training◐ Lab's own use◐ Public release
Enact it on the menuSee the findings

What it does

If testing shows a model crossing a danger line, such as giving real help with a bioweapon, the lab must have experts attack it hard and fix what they find before anyone uses it.

Applies to: Any frontier model whose evaluations cross a defined capability threshold.

Where things stand

Both major frameworks already have this shape. OpenAI's Preparedness Framework requires safeguards at High capability; its GPT-6 Astra card says safeguards were required even for internal deployment because of cyber capability. Anthropic's RSP has ASL-3 and ASL-4 safeguards tied to thresholds. Thresholds are now being crossed routinely rather than hypothetically.

Why it costs almost no lead

It only applies when a model crosses a danger line, and most fixes are prepared in advance. Averaged over a year, the effect is under a day.

Biggest unknown: How often thresholds get crossed and whether safeguards must be in place for internal use as well as release.

Why it lowers p(doom) by ~0.15%

When a model shows warning signs, outside experts try to break it before release.

It targets the riskiest models, but only if the warning signs are noticed.

The strongest case that it costs more

Thresholds only get crossed more often from here. If every frontier model now triggers elevated testing, the conditional becomes unconditional and this item becomes a standing four-to-six-week gate. Worse, if regulators rather than labs define the threshold and the remediation standard, the lab loses the ability to prepare safeguards in advance for a threshold it can anticipate. That would move the central price toward a month and push more of it onto the internal clock.

The debate

For

Sources

Rough starting points, not precise forecasts. Lead costs assume China doesn't depend on U.S. models, the case least favorable to safety laws, and count 3 years. On the menu you can change every assumption and put in your own numbers. Last priced 2026-09-26.