Detailed Analysis
Anthropic's practice of physically destroying millions of purchased print books after scanning them for AI training data sits at the center of a landmark copyright dispute that has significant implications for the entire AI industry. According to court filings in the case Bartz v. Anthropic, the company bought used books in bulk, cut off their bindings, scanned the pages into digital files, and then discarded the physical remains—all to build a training corpus for its Claude models. This practice came to light during litigation brought by authors who allege that Anthropic's use of their copyrighted works, both through this destructive scanning process and through the separate acquisition of pirated book datasets, violated their intellectual property rights.
The legal question is more nuanced than a simple yes-or-no on legality. In June 2025, federal judge William Alsup issued a mixed ruling that distinguished between the two methods Anthropic used to obtain training material. He found that Anthropic's practice of purchasing physical books, digitizing them, and destroying the originals likely qualified as fair use under U.S. copyright law, drawing an analogy to a consumer's right to format-shift media they have legally purchased—similar to ripping a CD to a digital library. However, Alsup was far less sympathetic to Anthropic's separate acquisition of millions of books from pirated sources like LibGen and Books3, ruling that this method was not protected by fair use because the company never paid for or legitimately acquired those copies in the first place. This distinction proved costly: Anthropic subsequently agreed to a $1.5 billion settlement, one of the largest copyright settlements in history, to resolve claims tied to the pirated works, while the "buy-scan-destroy" method for legitimately purchased books remained comparatively insulated from liability.
This case matters because it establishes some of the first significant judicial guidance on how copyright law applies to the immense data-hunger of large language models. AI companies including Anthropic, OpenAI, Meta, and Google have all faced lawsuits alleging that their models were trained on copyrighted books, articles, and other creative works without permission or compensation. The ruling suggests that courts may be willing to extend fair use protections to companies that at least legally purchase source material before digitizing it, even if the ultimate use—training a commercial AI product—is transformative in ways authors never anticipated or consented to. At the same time, the harsh treatment of pirated sourcing signals that courts will not extend similar leniency to companies that skip the step of legitimate acquisition, regardless of how the data is subsequently used.
Beyond the legal technicalities, the story reflects broader tensions in the AI industry around data provenance, transparency, and the economic value of creative labor. Authors and publishers have increasingly organized to challenge what they see as wholesale appropriation of their work by well-funded AI labs, while those companies argue that access to vast bodies of human writing is essential to building capable language models. The $1.5 billion settlement figure signals that regulators, courts, and rights holders are beginning to impose real financial consequences on AI developers for data-sourcing practices, potentially reshaping how the industry acquires training data going forward—pushing companies toward licensing deals, provable legitimate acquisition, and more transparent data pipelines rather than opportunistic scraping or reliance on pirated repositories.
Read original article →