In August 2025, Anthropic settled a class action lawsuit covering roughly 500,000 works for $1.5 billion — the largest copyright recovery in US history. Two months earlier, a federal judge had found that the same company's use of books to train Claude was "transformative — spectacularly so," suggesting the fair use argument could prevail in some contexts. These two outcomes are not contradictory. They reflect how genuinely unsettled the legal terrain is, and how high the stakes have become for everyone involved.
The argument: the AI copyright dispute is not resolvable through optimization. There is a real trade-off between two legitimate interests — the creative ecosystem's ability to sustain itself and the development of AI systems that depend on large corpora of human expression — and courts, legislatures, and the industry are collectively being forced to decide how to allocate it. The decision will have lasting consequences for both sides, and the framing matters as much as the outcome.
What the courts actually decided in 2025
Three federal district court decisions in 2025 began to sketch the doctrine. IPWatchdog's year-end review provides the clearest summary of where each case landed.
The Bartz v. Anthropic decision turned on transformativeness. Judge Alsup found that training a language model on copyrighted books constitutes fair use at the ingestion stage because the use is transformative — the model doesn't reproduce books, it learns statistical patterns from them. This is the strongest version of the fair use argument, and it's contested. Critics argue that the line between "learning patterns" and "memorizing text" is not as clear as the framing implies; researchers have demonstrated that frontier models can reproduce substantial passages verbatim under the right prompting conditions.
Germany's GEMA v. OpenAI reached a different conclusion on similar facts. The Hamburg court found that training GPT-4 and GPT-4o on copyrighted song lyrics infringed German copyright — transformativeness doctrine doesn't have the same weight under EU law, and Germany's approach treats the rights-holder's ability to license as primary. The practical implication is that AI companies training on European-hosted data face different legal exposure than training on data hosted in the US.
The Morrison Foerster 2026 outlook suggests litigation may shift in 2026 from training data to AI outputs — not whether models were trained on copyrighted material, but whether outputs constitute derivative works of their training data. This is the harder doctrinal question and potentially the more consequential one for how AI systems are designed and deployed.
The underlying trade-off
There are two interest groups in this dispute who both have reasonable claims, and the interesting question is how to structure the trade-off between them rather than which side is right.
The AI developers' position is that large-scale training on the corpus of human knowledge is foundational to building useful AI systems, that transformative use of information to enable learning is protected under fair use doctrine and its international analogues, and that requiring licensing for every training document would be administratively unworkable and economically prohibitive. The US Copyright Office's Part 3 report on generative AI training takes this position seriously and stops short of recommending a per-use licensing requirement.
The creative community's position is that AI systems are trained on creative work that has market value, that this training benefits AI companies commercially without compensating the people who produced it, and that the competitive pressure this creates threatens the economic viability of creative careers. The argument isn't that AI learning from human writing is categorically wrong — it's that the current arrangement transfers value from creators to AI developers without any mechanism for sharing the gains. The $1.5 billion settlement suggests this argument has enough legal legs to make litigation an attractive alternative to negotiation.
Both positions are internally coherent. The disagreement is not about facts but about which legal framework should govern a novel situation.
The transparency question
One development that cuts across the litigation debate is the push for training data transparency. Research published in the Journal of Intellectual Property Law & Practice argues that transparency requirements — mandating that AI companies disclose what data their models were trained on — could resolve some of the uncertainty without requiring courts to decide the underlying fair use question.
The logic: if AI companies disclose training data composition, rights-holders can identify whether their work was included and make informed decisions about whether to seek licensing agreements, participate in opt-out registries, or litigate. This creates a market for training data licensing rather than forcing the question to court.
The counter-argument is that training data disclosure is commercially sensitive and technically complex. Pretraining corpora for frontier models contain billions of documents from thousands of sources; documenting them in a way that's useful to rights-holders is a significant engineering and legal challenge. Labs also argue that training data composition is a competitive advantage.
This tension — between transparency as a solution and transparency as a cost — is likely to be the site of the next regulatory battle, particularly in the EU where the AI Act's training data disclosure requirements are still being interpreted.
The creative ecosystem argument
The claim that unlicensed AI training threatens the creative ecosystem deserves scrutiny rather than assumption.
The strong version of the argument is that AI systems trained on creative work will produce outputs that substitute for human creative work at scale, depressing demand for original human creation, which reduces the incentive to produce creative work, which shrinks the corpus available for future training. This is a real dynamic — but it depends on AI outputs being close substitutes for human creative work in ways that aren't universally true.
The weaker version — that the current arrangement transfers value from creators to AI companies without compensation — is more clearly supported. Even if AI-generated content doesn't fully substitute for human creation, the fact that AI capabilities were built on a corpus of human-created work without licensing revenue flowing to creators is an economically significant redistribution.
The European Parliament study on generative AI and copyright attempts to quantify this and finds the amounts are large — but faces the fundamental measurement problem that the counterfactual (what would AI development have looked like with mandatory licensing) is inherently speculative.
What different resolution mechanisms would produce
Full fair use protection for training data means AI companies can continue training on available internet text without licensing requirements. This maximizes AI development speed and minimizes costs. It also means the creative ecosystem receives nothing for its contribution to the training corpus, and future creators face a harder economic environment if AI outputs erode demand for human-created content.
Mandatory licensing per training document would make frontier model training prohibitively expensive in its current form. This might produce smaller models trained on licensed datasets, more data-efficient training methods, or concentration of AI development among entities that can negotiate bulk licensing. It wouldn't necessarily stop AI development, but it would restructure it significantly.
Opt-out registries — where creators can signal that their work should not be used for training — are the current practical middle ground. The Copyright.gov report discusses this approach. The criticism is that the default-in structure places the burden on creators who would prefer to control their work's use, rather than requiring AI companies to affirmatively license before training.
Collective licensing schemes — where a rights-management organization negotiates on behalf of creators and distributes licensing revenue proportional to use — are the European default for related rights disputes (radio airplay, etc.) and have been proposed for AI training. The challenge is the same as for transparency: knowing which works were used and at what frequency is technically complex, and the administrative costs of distribution at scale are significant.
The question the litigation doesn't answer
Courts are deciding whether specific uses of copyrighted material for AI training constitute fair use under existing doctrine. This is an important question and the answer matters. It is not the same question as: what policy regime produces the best long-run outcomes for creative production and AI development simultaneously?
The legal question has a coming answer — one that will be shaped by a handful of district court decisions, some appellate rulings, and eventually possibly Supreme Court review. The policy question — how to structure a relationship between AI development and the creative ecosystem that sustains both — will outlast any particular legal ruling.
Warning that requiring licensing would throttle transformative technology is probably true at current licensing cost structures. Warning that unlicensed training will corrode the creative ecosystem is probably also true in the long run if AI outputs substitute sufficiently for human creative work. The interesting design space is between these poles: mechanisms that allow AI development to continue while ensuring that creators receive some share of the value their work enables.
No court is positioned to design that. The litigation will produce a legal framework. Whether that framework is also a good policy is a different question — and one that regulators, legislators, and the industry itself will have to answer in the space the courts leave open.
References
- Copyright and AI Collide: Three Key Decisions from 2025 — IPWatchdog
- Copyright and Artificial Intelligence Part 3: Generative AI Training — US Copyright Office
- AI Trends for 2026: Copyright Litigation Shifts from Training to Outputs — Morrison Foerster
- Copyright and AI Training Data — Transparency to the Rescue? — Journal of IP Law & Practice
- Generative AI and Copyright: Training, Creation, Regulation — European Parliament (2025)
- Copyright & AI in the UK: The Debate Rolls On — Two Birds (2026)