Safety tests
Mandatory internal capability evaluations. Run dangerous-capability evaluations on every frontier model before you deploy it.
What it does
Before release, the lab must test every frontier model for dangerous abilities, such as helping build a bioweapon or break into computer systems, and record the results.
Applies to: Frontier models before external deployment.
- Before deployment, run a defined suite of evaluations for chemical, biological, cyber and autonomy capabilities and document the results.
- Compare results against published thresholds and record the deployment decision.
Where things stand
OpenAI's Preparedness Framework says every covered model undergoes its suite of scalable evaluations prior to deployment. Anthropic runs capability assessments under its RSP. GPT-5.5 went through targeted red-teaming for cyber and biology before release. The requirement is already the norm.
Why it costs almost no lead
Labs already run these tests, mostly on earlier versions of the model. Making them mandatory adds almost nothing to the lab's own progress; the public launch waits a few days.
Biggest unknown: How much of the evaluation must wait for the final checkpoint rather than a proxy.
Why it lowers p(doom) by ~0.15%
Catches dangerous abilities, like helping build weapons, before a model is released.
Testing is the main way labs find dangerous abilities. Labs already test, so a mandate mostly locks in and standardizes the practice.
The strongest case that it costs more
Evaluation suites grow. A mandate with a fixed list could add human uplift studies or agentic evaluations that take weeks and cannot run on proxies. And a legal duty to evaluate is a legal duty to act on the result, which turns a measurement into a gate. That gate is priced separately, but a critic would say the two cannot be separated in practice.
The debate
For
- Shevlane et al. researchers argued dangerous-capability evaluations should drive training and release decisions, 2023.
- EU Code of Practice requires state-of-the-art evaluations before a model is placed on the market, 2025.
- California SB 1047 would have required a critical-harm assessment before release, 2024.
Against
- Meta refused to sign the EU code, calling it overreach, 2025.
- Dean Ball, Foundation for American Innovation called a federal pre-release evaluation mandate a federally mandated veto point, 2025.
Sources
- AISI, Early lessons from evaluating frontier AI systems: light tier: few days, 1–2 weeks end to end; standard: 1–2 weeks; elevated: several weeks
- OpenAI Preparedness Framework v2: every covered model undergoes scalable evaluations prior to deployment
- GPT-5.5 System Card: full pre-deployment suite plus targeted red-teaming
Rough starting points, not precise forecasts. Lead costs assume China doesn't depend on U.S. models, the case least favorable to safety laws, and count 3 years. On the menu you can change every assumption and put in your own numbers. Last priced 2026-09-26.