Red team if risky
Triggered red-team and remediation. If an evaluation crosses a danger threshold, do elevated testing and fix it before anyone can use it.
What it does
If testing shows a model crossing a danger line, such as giving real help with a bioweapon, the lab must have experts attack it hard and fix what they find before anyone uses it.
Applies to: Any frontier model whose evaluations cross a defined capability threshold.
- When a pre-deployment evaluation crosses a capability threshold, conduct elevated testing (expert red-teaming, possibly uplift studies).
- Implement and verify safeguards before deployment; if the threshold governs internal use, before internal deployment too.
Where things stand
Both major frameworks already have this shape. OpenAI's Preparedness Framework requires safeguards at High capability; its GPT-6 Astra card says safeguards were required even for internal deployment because of cyber capability. Anthropic's RSP has ASL-3 and ASL-4 safeguards tied to thresholds. Thresholds are now being crossed routinely rather than hypothetically.
Why it costs almost no lead
It only applies when a model crosses a danger line, and most fixes are prepared in advance. Averaged over a year, the effect is under a day.
Biggest unknown: How often thresholds get crossed and whether safeguards must be in place for internal use as well as release.
Why it lowers p(doom) by ~0.15%
When a model shows warning signs, outside experts try to break it before release.
It targets the riskiest models, but only if the warning signs are noticed.
The strongest case that it costs more
Thresholds only get crossed more often from here. If every frontier model now triggers elevated testing, the conditional becomes unconditional and this item becomes a standing four-to-six-week gate. Worse, if regulators rather than labs define the threshold and the remediation standard, the lab loses the ability to prepare safeguards in advance for a threshold it can anticipate. That would move the central price toward a month and push more of it onto the internal clock.
The debate
For
Sources
- AISI, Early lessons from evaluating frontier AI systems: elevated tier: several weeks
- GPT-6 Astra deployment safety page: safeguards required even for internal deployment
- OpenAI Preparedness Framework v2
- Anthropic Responsible Scaling Policy
Rough starting points, not precise forecasts. Lead costs assume China doesn't depend on U.S. models, the case least favorable to safety laws, and count 3 years. On the menu you can change every assumption and put in your own numbers. Last priced 2026-09-26.