Limit AI-run research
Human sign-off on automated AI research. AI agents can assist with AI research, but humans must direct and approve it.
What it does
Labs may not let AI systems run AI research on their own from start to finish. A human researcher has to direct and approve the experiments AI agents run, and labs must report how much of their research AI is doing.
Applies to: Frontier developers using AI systems to automate their own research.
- Require a human to direct and approve research experiments carried out by AI agents inside frontier labs.
- Prohibit fully autonomous AI research loops above a capability threshold.
- Report the share of research work done by AI systems.
Where things stand
Both major frameworks already treat automated AI research as a danger sign. Anthropic's Responsible Scaling Policy v3.4 has an automated R&D threshold met if its models could fully substitute for its research scientists and engineers, or if AI progress doubled in speed. OpenAI's Preparedness Framework lists fully automated AI R&D as a Critical capability. Neither framework limits today's partial automation.
Why it costs ~12 days
Human sign-off caps how much AI can speed up research. AI's research boost is still modest, so the loss over a year is about 4 days.
Biggest unknown: How much faster AI-run research is making the labs, which they are only starting to measure.
Why it lowers p(doom) by ~0.25%
Keeps humans in charge as AI starts doing AI research, which could otherwise speed up beyond anyone's oversight.
Many researchers see AI automating AI research as the step where progress could outrun oversight. Keeping humans in the loop targets that directly.
The strongest case that it costs more
This is the policy most likely to get more expensive over time. If AI research automation is the main driver of progress in 2027 and beyond, keeping humans in every loop could halve the pace of the frontier. China's labs would face no such limit.
The debate
For
- IDAIS-Beijing scientists including Bengio said no AI should improve itself without explicit human approval, 2024.
- ControlAI, A Narrow Path would prohibit using AI systems to build or improve new ones, 2024.
- Google DeepMind added protocols for AI that could speed up AI research to destabilizing levels, 2025.
Against
- James Pethokoukis, AEI argued a superintelligence ban would freeze progress and hand China an advantage, 2025.
Sources
- Anthropic, When AI builds itself: 80% of merged code by Claude; engineers ship 8x as much code
- OpenAI, Research acceleration: the view inside OpenAI (Sept 2026): automated research intern goal met
- Anthropic Responsible Scaling Policy v3.4: automated R&D threshold
- OpenAI Preparedness Framework v2: AI self-improvement: Critical means fully automated AI R&D
Rough starting points, not precise forecasts. Lead costs assume China doesn't depend on U.S. models, the case least favorable to safety laws, and count 3 years. On the menu you can change every assumption and put in your own numbers. Last priced 2026-09-26.