Limit AI agents
Limits on autonomous AI agents. AI agents need a person's approval to move money, deploy code, get computing power or copy themselves.
What it does
Frontier AI agents may not take consequential actions on their own: moving money, deploying code to live systems, acquiring computing power or copying themselves. A person has to approve each one, and no agent above a capability threshold may run open-ended without human oversight.
Applies to: AI agents built on frontier models, including those labs use internally.
- AI agents need a person's approval before moving money, deploying code to live systems, acquiring computing power or copying themselves.
- No agent above a capability threshold may run open-ended without human oversight.
- Labs must log agent actions for audit.
Where things stand
Human confirmation before purchases or code changes exists only as voluntary product design, such as coding agents that ask before editing files. No law requires it.
Why it costs ~3 weeks
Agents are a growing share of what labs sell and of how they do research, so requiring sign-off slows both a little. About 6 days a year in.
Biggest unknown: How to define a consequential action, and whether sign-off stays meaningful when agents take thousands of actions a day.
Why it lowers p(doom) by ~0.2%
Keeps AI agents from moving money, grabbing computing power or copying themselves without a person approving.
Loss of control is most likely to start with an agent acting on its own: acquiring resources, copying itself or resisting shutdown. Requiring a human to approve those actions targets that path directly, though approvals can become a rubber stamp at scale.
The strongest case that it costs more
At the scale agents run, sign-off is either rubber-stamping or a brake that kills most of their value. 'Consequential action' is hard to define, and a U.S. rule would push use toward foreign and open-weight agents while slowing defenders who need machine-speed responses.
The debate
For
- Margaret Mitchell and colleagues at Hugging Face argued that fully autonomous AI agents should not be developed, because the risk to people rises with each level of autonomy, 2025.
- Stuart Russell and colleagues proposed verifiable red lines, including no AI copying or improving itself without human approval and no resisting shutdown, 2025.
- Yoshua Bengio and colleagues argued that generalist AI agents pose catastrophic risks including loss of control, and proposed non-agentic 'Scientist AI' as a safer path, 2025.
Against
- Anthropic argued that agents must be able to work autonomously to be valuable, favoring calibrated oversight over blanket approval while still asking for sign-off before high-stakes actions, 2025.
- Security researchers at Black Hat argued that human sign-off doesn't scale to thousands of agents and risks becoming theater, favoring permissions and monitoring instead, 2026.
- Matt Perault, a16z argued that policy should regulate harmful uses of AI rather than AI development, 2025.
Sources
Rough starting points, not precise forecasts. Lead costs assume China doesn't depend on U.S. models, the case least favorable to safety laws, and count 3 years. On the menu you can change every assumption and put in your own numbers. Last priced 2026-09-26.