Detailed Analysis
Anthropic's practice of physically destroying millions of print books after scanning them for AI training data emerged as a central and controversial detail in the landmark copyright case Bartz v. Anthropic. Court filings revealed that the company purchased used physical books in bulk, stripped their bindings, scanned each page, and then discarded the paper originals—keeping only digital copies to train its Claude language models. This practice, while startling when framed as "book destruction" or "book burning," was actually a calculated legal strategy: Anthropic's lawyers argued that buying a physical book and converting it to a digital format for internal use constituted a lawful, transformative act under fair use doctrine, distinct from simply downloading pirated copies from shadow libraries like LibGen or Books3.
The distinction mattered enormously in court. In June 2025, U.S. District Judge William Alsup issued a split ruling: he found that Anthropic's practice of buying physical books, destroying them, and digitizing the content for training purposes was likely protected under fair use, comparing it to a form of format-shifting akin to what courts have allowed for personal use of purchased media. However, Alsup ruled that Anthropic's separate practice of downloading millions of pirated books from illegal online repositories—rather than purchasing and destroying physical copies—was not fair use and exposed the company to massive liability. This led to Anthropic's settlement of the case in September 2025 for $1.5 billion, one of the largest copyright settlements in history, covering an estimated 500,000 or more works whose authors could claim damages for pirated use.
The book destruction detail matters because it illustrates how AI companies are contorting themselves to find legally defensible pathways to the enormous troves of text needed to train large language models. High-quality, long-form writing found in books is considered exceptionally valuable training data compared to scraped web content, because books tend to be well-edited, structurally coherent, and rich in narrative and reasoning patterns. Rather than risk further piracy liability, Anthropic apparently calculated that buying millions of secondhand books and physically destroying them after scanning was a legally safer, though ethically and environmentally strange, alternative—raising uncomfortable questions about the human and cultural cost of feeding AI's data hunger, even when done "legally."
More broadly, this episode reflects the scramble across the AI industry to secure legitimate data sources as courts, regulators, and publishers increasingly scrutinize how models are trained. OpenAI, Meta, and Google all face similar lawsuits alleging unauthorized use of copyrighted books and articles, and the outcomes of these cases are reshaping industry norms around data licensing. Anthropic's $1.5 billion settlement and its unusual book-destruction strategy signal that the era of quietly scraping the internet and shadow libraries for training data is ending, replaced by a more litigious and cost-intensive landscape where companies must either license content outright or engineer creative, if bizarre, workarounds—like buying and destroying millions of physical books—to claim clean legal title to the data powering their models.
Read original article →