4-week test period
Twenty business days of evaluator access. Outside evaluators get at least four weeks with the model before it ships.
What it does
Outside testers get at least four weeks with a new model before the public can use it. Today they often get between three days and three weeks.
Applies to: Frontier models before external deployment (the EU Code of Practice's suggested period).
- Provide independent external evaluators with at least 20 business days of access to the model before public release.
- The EU Code of Practice describes this as appropriate for most systemic risks and evaluation methods; METR's guidance for lab staff repeats it.
Where things stand
Actual windows in 2025 and 2026 ran from three days (GPT-6 Astra, Apollo) to three weeks (o3, GPT-5). Ten business days on Claude Opus 5.5 is the recent norm at the generous end. A 20-business-day floor is two to three weeks more than typical practice.
Why it costs almost no lead
Only the public launch waits. The lab keeps building on the model internally, so the lead barely moves. If China depends on U.S. models, the delay slows China slightly more than America.
Biggest unknown: Whether labs respond by sharing checkpoints four weeks earlier or by shipping four weeks later.
Why it lowers p(doom) by ~0.15%
Gives government testers time to find dangers before the public gets a model.
Extra time helps only with dangers testers can find in a few weeks, and it doesn't cover the lab's own use of the model.
The strongest case that it costs more
Four weeks is a real number in a field where METR measures the frontier doubling every 105 days. If commercial availability drives revenue, and revenue drives compute purchases, a systematic four-week delay on every release compounds into a slower frontier. And the internal-versus-public distinction may be thinner than it looks: OpenAI's internal-to-external gap is only a few weeks, so a four-week floor could exceed the lab's own cadence and force it to hold the internal frontier too.
The debate
For
- EU Code of Practice says at least 20 business days is appropriate for most evaluations, 2025.
- METR said its o3 evaluation was done in a relatively short time, over only three weeks, 2025.
- Charnock, Casper and colleagues noted outside testers got about one week with Claude Sonnet 4 and 4.5 and argued for more, 2026.
Sources
- EU GPAI Code of Practice, Safety and Security chapter, Appendix 3.4: at least 20 business days is appropriate for most systemic risks
- METR, Frontier AI safety regulations: a reference for lab staff
- METR evaluation of Claude Opus 5.5: 10 business days
- GPT-6 Astra deployment safety page: 3 days for Apollo
- METR Frontier Risk Report PDF: OpenAI's internal-to-external gap was a few weeks
Rough starting points, not precise forecasts. Lead costs assume China doesn't depend on U.S. models, the case least favorable to safety laws, and count 3 years. On the menu you can change every assumption and put in your own numbers. Last priced 2026-09-26.