AI Training and Copyright: Where Does the Law Stand?

AI Training and Copyright: Where Does the Law Stand?

Listen to the Article

Generative AI has raised an uncomfortable question for copyright law: When a model learns from published works, is that lawful use or infringement?

Courts across the US, EU, and UK are actively trying to reach a verdict, and the stakes are significant. After all, AI developers, publishers, and rights holders are all waiting on rulings that could reshape how content is created, licensed, and protected. This article examines the legal concepts at play, the practical challenges they raise for legal professionals, and what recent case outcomes suggest about the future of AI training.

Piracy vs. Fair Use: The Crucial Distinction for AI

The debate often gets framed as a simple binary: AI training is either piracy or it is not. However, the legal considerations are much more nuanced than that.In terms of copyright, piracy refers to the unauthorized reproduction or distribution of a protected work, typically for direct personal or commercial consumption. Common examples include downloading pirated films or distributing bootleg software, which is conducted with no plausible transformative purpose and clear market impact.But AI training is inherently different. When a large language model (LLM) is trained, it processes copyrighted text, images, or audio to extract statistical patterns in how words relate or visual elements combine.These are then converted into mathematical representations, but the original work is not stored or distributed as a final output, though intermediate copies may be made during data collection and preprocessing. The ingested materials become a whole new model capable of generating content, not a reproduction of any specific work.One of the cornerstones of US copyright law is the doctrine of fair use. To assess whether a case constitutes fair use, courts routinely assess four statutory factors, including whether the use is transformative and adds new meaning or purpose.

Copyrighted works are not consumed as creative expressions, but transformed into statistical data over the course of AI training. To put it differently, these works are used to build something fundamentally different, not a substitute that would (typically) compete in the same market.The EU takes a different approach, relying on two statutory text and data mining (TDM) exceptions introduced by the Copyright in the Digital Single Market Directive (2019):

  • Article 3 permits research organizations to mine works they have lawful access to, with no opt-out available to rights holders

  • Article 4 permits commercial TDM on lawfully accessed works, but rights holders may expressly reserve their rights to exclude this use. 

The UK’s TDM exception is currently limited to non-commercial research purposes.Applicable law often considers the physical location of the training process, but courts may also apply choice-of-law rules and look to where data was acquired, accessed, or distributed.

The Challenges Facing Legal Practitioners

For copyright lawyers and in-house legal teams, this evolving landscape presents several concrete challenges.The law is unsettled and will remain so for some time. As of mid-2026, there is no settled appellate authority on liability for generative-AI training in the US. Three summary judgments on fair use have been issued at the District Court level, in Thomson Reuters v. Ross Intelligence, Bartz v. Anthropic, and Kadrey v. Meta, with conflicting reasoning across decisions that varied in procedural scope. For this reason, advising clients in this environment requires carefully qualifying every position taken.The fair use analysis is intensely fact-specific. The aforementioned summary judgments illustrate how much the outcome can vary depending on the model’s purpose. In Thomson Reuters v. Ross Intelligence, the fair use defense failed because the AI tool, trained on Westlaw headnotes, was not sufficiently transformative, as it did not have a “further purpose or different character.” 

In Bartz v. Anthropic and Kadrey v. Meta, the defense succeeded, with courts describing the training of general-purpose generative AI models as “highly,” “exceedingly,” and “quintessentially” transformative. 

One significant variable: how broadly the model is designed to function. Although courts weigh all four fair-use factors, more targeted AI tools appear to face greater scrutiny than general-purpose ones.Market harm is a heavily weighted factor, but also the hardest to prove. Judge Chhabria (Kadrey v. Meta case) identified market dilution as potentially the strongest argument for plaintiffs, as AI outputs could flood the market with works similar to those used in training, depressing their value. However, the plaintiffs presented no evidence on this point, and the argument failed as a result. Future litigation is likely to be fought on this ground, requiring expert economic evidence and detailed analysis of downstream model outputs.The EU opt-out mechanism creates compliance uncertainty on both sides. Rights holders must “expressly reserve” their rights in a machine-readable format, but no legally mandated standard exists. National courts in some EU member states have questioned whether natural-language reservations, such as those embedded in a website’s privacy policy, meet the machine-readable standard. Effective opt-outs may need to be machine-readable in a format that automated crawlers can detect and process, though the precise technical standard remains legally unsettled. For legal advisers working with content businesses or AI developers in the EU, this is an active compliance risk that they need to be aware of.Training data provenance is increasingly material. Both the Bartz and Kadrey decisions engaged directly with whether AI companies sourced training data from pirated “shadow libraries” or online repositories of unlicensed books. In cases where training data is sourced unlawfully, it can affect the fair use analysis and, independently, expose developers to direct infringement liability for the act of acquisition, even if the subsequent training would otherwise qualify as fair use.

The Anthropic Settlement: A Signal Worth Reading

The most consequential development to date is the $1.5 billion settlement, which was granted final approval in July 2026. The plaintiff counsel in Bartz v. Anthropic has described the instance as the “largest known copyright recovery in history.”The class action lawsuit was brought forward by over 300,000 writers who argued that Anthropic had used their works without permission to train its Claude AI system. The court found that AI training on books could, in principle, qualify as fair use. However, it also found that Anthropic had obtained a significant portion of its training data through pirated book repositories, instead of legitimate channels. That finding created distinct liability, separate from the fair use question.The settlement covers more than 482,000 books and provides payments of approximately $3,000 per title, with around 91% of covered works already claimed as of July 22.Anthropic acknowledged the outcome but emphasized the court’s fair use finding as the more significant legal point.What are the main lessons here?For starters, the source of training data matters as much as how it is used. Even where the training process itself is defensible as fair use, unlawfully obtained data creates independent liability. Legal teams advising AI developers should treat data provenance as a core due diligence item, not an afterthought.Settlements of this scale are likely to shape litigation strategy across the industry. Rights holders now have a clearer sense of what claims are worth pursuing and on what basis. AI developers, for their part, have a stronger incentive to establish legitimate data acquisition processes from the outset.Finally, the Anthropic case does not close the fair use debate, but simply narrows one part of it. The interaction between pirated source material and fair use remains contested. In Kadrey v. Meta, Judge Chhabria declined to treat shadow library sourcing as an automatic loss for the defendant, noting that the plaintiffs had not submitted relevant evidence.The two decisions reach different outcomes on shadow library sourcing, though on different evidentiary records, which appellate courts will eventually need to address.

AI Copyright Law Is Still Shifting

The legal framework governing AI training is still in motion.For legal professionals, the priority is clarity about what remains unresolved. Fair use analysis depends on facts: the model’s purpose, its outputs, the source of its training data, and the evidence of market harm. Advising clients, whether they’re rights holders or AI developers, requires understanding the variables and nuances, rather than applying fixed rules.In the EU, opt-out compliance is an immediate practical concern. Meanwhile, in the US, litigation strategy around market dilution is emerging as the contested frontier. Ultimately, staying vigilant and up to speed with the case developments is the only way forward.

Subscribe to our weekly news digest.

Join now and become a part of our fast-growing community.

Invalid Email Address
Thanks for Subscribing!
We'll be sending you our best soon!
Something went wrong, please try again later