Pay for training data
Consent to train on copyrighted work. No training on copyrighted works or personal data without clear, affirmative consent.
What it does
AI companies must get clear permission, usually a paid license, before training on copyrighted books, articles, art or personal data. Anyone whose work is used without consent can sue.
Applies to: Training of generative AI systems on copyrighted works or personal data.
- Obtain clear, affirmative consent before using copyrighted works or personal data, including to train a generative AI system (AI Accountability and Personal Data Protection Act, S. 2367).
- Give individuals a right to sue for use without consent.
Where things stand
Labs have begun licensing some data from publishers and paid to settle past claims. Anthropic's $1.5 billion settlement with authors covered about 482,000 pirated books and only past conduct. Most training data is still used without explicit consent.
Why it costs ~3 months
Removing unlicensed data means training on less, or repeated, data. The model estimates about a month of lost progress a year in.
Biggest unknown: How much training data labs could license quickly, and how much capability depends on the rest.
Its effect on p(doom)
Protects writers, artists and other creators whose work trains AI.
A copyright and fairness rule. It doesn't target catastrophic risk.
The strongest case that it costs more
Chinese labs would not follow U.S. consent rules. If the rule removes a large share of high-quality text and code, the capability hit could persist for years, not a single training run.
The debate
For
- Sens. Hawley and Blumenthal introduced a bill barring training on copyrighted works without consent, 2025.
- Authors Guild welcomed the bill, 2025.
- News/Media Alliance backs licensing of news content for AI training, though not a mandate, 2025.
Against
- Google and OpenAI told the White House fair use is critical to AI development, 2025.
- Chamber of Progress called the Hawley–Blumenthal bill a direct attack on fair use, 2025.
Sources
- S. 2367 bill text (govinfo): covers training of a generative AI system
- Hawley and Blumenthal unveil bipartisan bill (July 2025)
- Authors Alliance, Bartz v. Anthropic settlement gets preliminary approval: $1.5B for about 482,460 books
- Epoch AI, Will we run out of data?: public text stock used up between 2026 and 2032
Rough starting points, not precise forecasts. Lead costs assume China doesn't depend on U.S. models, the case least favorable to safety laws, and count 3 years. On the menu you can change every assumption and put in your own numbers. Last priced 2026-09-26.