TECH NEWS

Is It Legal to Train AI Models on Copyrighted Books? The Complicated Truth in 2026

The explosive growth of generative artificial intelligence has sparked one of the most significant copyright battles of the digital age. At the center of the fight sits a deceptively simple question: Can AI companies train their large language models on copyrighted books without the authors’ permission? As of late 2026, the answer emerging from U.S. courts is nuanced, carefully limited, and still evolving. Training itself is increasingly viewed as fair use when done properly. Building permanent libraries of pirated books is not.

The Fair Use Doctrine at the Heart of the Dispute

Under U.S. copyright law, the exclusive rights of authors are not absolute. Section 107 of the Copyright Act creates a flexible exception known as fair use. Courts examine four factors: the purpose and character of the use (especially whether it is transformative), the nature of the copyrighted work, the amount used, and the effect on the potential market for the original.

AI companies argue that training a large language model is profoundly transformative. The model does not store or reproduce books for readers. Instead, it analyzes statistical patterns across millions of texts so it can generate new sentences, answer questions, and produce original content. In this view, the process resembles a human writer reading widely and then creating something different. Authors and publishers counter that the scale is unprecedented, the commercial purpose is clear, and the resulting AI systems threaten to flood markets with competing works that undermine the economic foundation of professional writing.

Until 2025, these arguments remained largely theoretical. Then two federal judges in the Northern District of California issued the first major rulings addressing books specifically.

The Landmark 2025 Rulings

In June 2025, Judge William Alsup decided Bartz v. Anthropic. Authors Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson sued Anthropic, the company behind the Claude family of models, alleging that their books had been used without permission. The evidence showed Anthropic had obtained books in two distinct ways. Some physical copies were purchased, scanned, and digitized. Others were downloaded from notorious shadow libraries such as Books3, LibGen, and Pirate Library Mirror.

Judge Alsup drew a sharp distinction. Using lawfully acquired books to train the models, he ruled, was “quintessentially transformative.” The purpose was not to replicate or replace the original works but to enable the creation of something new. He compared the process to a student reading literature in order to become a better writer. Digitization of purchased print copies for internal training use was also fair use. However, downloading and retaining millions of pirated books to create a permanent central library was not. That conduct, the judge held, was “inherently, irredeemably infringing,” even if the copies were later used for training. Anthropic ultimately settled the remaining claims for $1.5 billion, the largest copyright settlement in U.S. history.

Two days later, Judge Vince Chhabria reached a similar bottom-line result in Kadrey v. Meta Platforms. Authors alleged that Meta had trained its Llama models on their books, including material from shadow libraries. The court found the training highly transformative and granted summary judgment on fair use grounds. Judge Chhabria placed heavier emphasis on the market-harm factor and noted that the plaintiffs had not presented strong evidence of market dilution. He also signaled that better-developed records in future cases could lead to different outcomes.

Together, the two decisions established an important early principle: the act of training a general-purpose generative model on copyrighted books can qualify as fair use, particularly when the works are lawfully obtained. The decisions also highlighted that acquisition methods and retention practices remain high-risk areas.

Why Piracy Changes Everything

The distinction between lawful and unlawful sourcing has proven decisive. Courts have been far more skeptical when companies rely on pirated datasets. Judge Alsup rejected the idea that a transformative purpose could cleanse the initial act of downloading unauthorized copies when legal alternatives existed. Creating a permanent, general-purpose library of stolen books, he ruled, stands outside the protection of fair use.

This ruling carries practical consequences. Many early large language models were trained, at least in part, on datasets assembled from shadow libraries. Companies that continue to retain those materials face ongoing exposure. Those that purchase books, license content, or carefully curate publicly available material operate on safer ground.

Broader Context and Remaining Uncertainties

These California decisions are not the final word. They are district-court rulings, fact-specific, and still subject to appeal. Other cases have produced different results when the AI system more directly competed with the original works. In Thomson Reuters v. Ross Intelligence, a court rejected fair use for an AI legal research tool that closely mirrored Westlaw headnotes, finding the use insufficiently transformative.

The U.S. Copyright Office has also expressed caution. In its reports on generative AI, the Office has noted that commercial training on vast quantities of expressive works raises serious questions under the market-harm factor and cannot be presumed fair use in every circumstance. Output-level infringement remains a separate issue. If a model produces text that is substantially similar to a protected work, authors can still pursue claims based on the generated content itself.

Outside the United States the legal landscape is often stricter. The European Union provides limited text-and-data-mining exceptions, particularly for research, but commercial training frequently requires licenses or respect for opt-outs. Many other jurisdictions lack a broad fair-use doctrine equivalent to the American one, making unauthorized commercial training more vulnerable to infringement claims.

Practical Implications for Creators and Companies

For authors, the current case law offers partial reassurance rather than complete protection. Training on their books may be lawful, yet the mass retention of pirated copies is not, and close reproduction in model outputs can still trigger liability. Some writers and publishers have begun exploring collective licensing arrangements as a pragmatic path forward.

For AI developers, the message is clearer on process than on principle. Lawful acquisition of training materials, limited retention tied to actual training needs, and technical safeguards against verbatim memorization reduce legal risk. Companies that treat shadow libraries as convenient free resources continue to invite substantial exposure, as Anthropic’s settlement illustrates.

The technology continues to advance faster than the law. Models grow more capable, training datasets expand, and new lawsuits test the boundaries established in 2025. Questions about market dilution—whether AI-generated books, articles, and stories will erode demand for human-authored works—remain largely unresolved and may prove decisive in future cases.

As of August 2026, the prevailing view in the most relevant U.S. courts is that training large language models on copyrighted books can constitute fair use when the books are obtained lawfully and the use is genuinely transformative. The same courts have drawn a firm line against the construction and long-term retention of pirated libraries. That distinction between transformative learning and simple theft of content is likely to shape the next phase of litigation and industry practice.

The debate is far from over. Appellate courts, additional trials, possible legislation, and evolving commercial realities will continue to refine the rules. For now, the law has begun to answer the central question with a carefully qualified yes: training can be legal, but the path taken to gather the books matters a great deal. Authors, technologists, and policymakers will spend the coming years testing just how far that qualified permission extends.

Click to rate this post!
[Total: 0 Average: 0]

About The Author

Leave a Reply

Discover more from NEWS NEST

Subscribe now to keep reading and get access to the full archive.

Continue reading

Verified by MonsterInsights