← Google News

Anthropic Destroyed Millions of Books to Train Claude: Was That Legal? - Yahoo

Google News · July 29, 2026
Anthropic Destroyed Millions of Books to Train Claude: Was That Legal? Yahoo [truncated: Google News RSS provides only a snippet, not full article

Detailed Analysis

Anthropic's practice of physically destroying millions of print books after scanning them for use in training Claude has become one of the more visually striking and legally consequential revelations to emerge from the wave of AI copyright litigation. According to court filings in the case brought by a group of authors, Anthropic purchased vast quantities of used books, cut off their bindings, and scanned the pages into digital format so the text could be fed into its large language models—destroying the physical originals in the process. The company has defended this approach as a legitimate way of establishing ownership of a lawfully acquired copy before digitizing it, distinguishing it from simply scraping pirated text off the internet. That distinction became central to a landmark ruling by U.S. District Judge William Alsup, who found that training an AI model on legally purchased books constitutes fair use, while separately ruling that Anthropic's earlier practice of downloading pirated books from shadow libraries did not enjoy the same protection.

The legal significance of this bifurcated ruling is substantial. Alsup's decision effectively created a roadmap for AI companies: acquiring content through legitimate purchase or license, even if it involves destructive digitization, can fall within fair use protections when the resulting copies are used for transformative purposes like training a model, whereas obtaining copyrighted material through piracy remains legally exposed regardless of downstream use. This nuance matters enormously for an industry that has been racing to amass training data at a scale that traditional licensing markets were never built to support. Anthropic still faces potentially massive damages tied to the piracy-sourced portion of its training corpus—reports have pointed to a settlement in the hundreds of millions of dollars range to resolve claims from authors whose works were allegedly pirated—even as the "legally purchased and destroyed" method has been validated as an acceptable, if unusual, workaround.

The book-destruction strategy also underscores how AI companies are contorting traditional intellectual property frameworks, built around physical scarcity and the first-sale doctrine, to fit an era where the economically valuable asset is not the object itself but the information it contains. By destroying the physical book after scanning it, Anthropic could argue it was not creating unauthorized additional copies in circulation—effectively converting one lawfully owned copy into one digital copy rather than multiplying the work. Whether this satisfies the spirit as well as the letter of copyright law remains contested among legal scholars, authors' groups, and publishers, many of whom argue that mass industrial-scale scanning for AI training purposes was never contemplated by fair use doctrine and represents a fundamentally different kind of use than, say, a library preserving a single copy for archival purposes.

More broadly, this episode reflects the collision between the AI industry's voracious appetite for training data and an intellectual property system designed for a pre-AI economy. Anthropic, OpenAI, Meta, and other major AI developers have all faced lawsuits alleging improper use of copyrighted books, journalism, and other creative works, and the outcomes of these cases are actively shaping how the entire industry sources data going forward. The Anthropic ruling, with its clear line between purchased-then-scanned material and pirated material, is likely to be cited repeatedly as other courts grapple with similar disputes, and it may push AI labs toward more elaborate—if legally defensible—physical acquisition strategies like bulk book buying and destruction, rather than risk the heavier liability associated with piracy. As authors, publishers, and courts continue to negotiate the boundaries of fair use in the generative AI era, this case stands as an early and unusually tangible example of how copyright law is being reinterpreted, quite literally book by book, to accommodate the mechanics of machine learning.

Read original article →