At a glance: The legal controversy is not simply whether an AI system can read a web page. Courts and publishers are contesting copying, training, retained datasets, generated outputs and commercial licensing under different legal rules.

Why the same phrase—fair use—produces different results
US fair use is a case-specific defence that weighs the purpose of use, nature of the work, amount used and effect on the market. It is not a blanket permission for every AI training activity. A court may view a particular training use as transformative while still treating the acquisition or retention of pirated source copies differently. The source of a dataset and what the model later outputs can matter.
The US Copyright Office’s case database lists both the 2025 Bartz v Anthropic and Kadrey v Meta decisions, with distinct outcomes and reasoning. They should not be generalized into a worldwide rule. Courts elsewhere apply different statutory exceptions and copyright frameworks.
What happened in a major author case
In July 2026, a federal judge approved Anthropic’s $1.5 billion settlement with authors over allegations involving pirated books used in model training. A settlement resolves the participating parties’ claims under its terms; it is not a universal Supreme Court ruling about all AI training or future publishing contracts.
Publishers are pressing their own cases
On October 8, 2026, USA Today Co and affiliated publications filed a new lawsuit alleging unauthorized use of their articles in training. Those are allegations in litigation, not established factual findings against the defendant. The broader landscape includes continuing disputes over newspaper, book and other protected creative material.
Robots.txt, paywalls and legal permission are different layers
| Mechanism | What it does | What it does not guarantee |
|---|---|---|
| robots.txt | Communicates crawling preferences to compliant bots. | Physically blocks all requests or resolves copyright law. |
| Authentication or paywall | Restricts access when correctly implemented. | Determines whether every otherwise lawful use is prohibited. |
| Terms and licensing | Creates contractual permissions and obligations where enforceable. | Automatically binds every actor in every jurisdiction. |
| Technical rate limits | Reduces abuse and server load. | Decides whether model training is fair use. |
The economic issue behind citations and royalties
Newsrooms fund reporting, editing and legal review, and argue that uncompensated training or competing answers can reduce their subscription and advertising markets. AI developers argue that training can be transformative and that models create new utility. Licensing arrangements offer one path to negotiated use, attribution and payment, but they raise difficult questions: which uses are covered, how value is allocated, and whether smaller publishers can bargain on fair terms.
A practical publisher checklist
- Audit which bots access the site and distinguish search indexing bots from AI training crawlers.
- Decide which material is intended for public discovery and which requires authenticated access.
- Publish clear licensing contacts, permissions and evidence of original publication dates.
- Preserve logs and document suspected infringement carefully; do not treat an automated bot label as proof of a violation.
- Review contracts and rights with counsel before making legal allegations or signing blanket licences.
What readers should take away
As of October 2026, there is no single legal rule that settles every AI training dispute across jurisdictions. The strongest reporting differentiates court holdings, settlements, complaints and industry opinions. Publisher rights, research access and competition are real interests that should be described without equating an allegation with a judgment.
Related on BCC: A separate case involving digital publishing and competition.
Related on BCC: Copyright case analysis.
Frequently asked questions
Does robots.txt enforce copyright law?
No. It is a crawler instruction convention, not a court order or a technical firewall.
Did the Anthropic settlement make all AI training illegal?
No. A case-specific settlement does not create that broad rule.
