← Google News

Why is Anthropic destroying books? | Kathryn James - The Guardian

Google News · August 5, 2026

Detailed Analysis

Anthropic's practice of physically destroying printed books to build a training dataset for its Claude models sits at the center of Kathryn James's Guardian piece, and the underlying facts trace back to revelations that emerged during the landmark copyright litigation Bartz v. Anthropic. Court filings showed that rather than relying solely on pirated digital text, Anthropic purchased millions of used print books, then had workers cut off their spines and run the loose pages through industrial scanners to convert them into machine-readable text. Once scanned, the physical books—often out-of-print, rare, or otherwise irreplaceable copies—were discarded. Judge William Alsup, who presided over the case, treated this practice favorably in his June 2025 ruling, finding that scanning legally purchased books for internal AI training constituted transformative fair use, even as he found Anthropic's separate use of pirated ebooks downloaded from shadow libraries like LibGen to be unlawful.

James approaches the story not as a technology reporter but as a rare-book specialist—she is associated with Yale's Beinecke Rare Book and Manuscript Library—and her framing foregrounds a dimension largely absent from the legal and business coverage of the Anthropic case: books as physical, historical objects rather than interchangeable vessels for text. From this vantage point, the practice of destructively scanning books raises questions that go beyond copyright economics. Print editions carry marginalia, printing variants, bindings, provenance markings, and other material evidence that scholars use to understand how texts were produced, circulated, and read. Treating a physical book as disposable once its "content" has been extracted reflects a view of literature as pure information, stripped of the artifact's cultural and historical value—a perspective increasingly at odds with how libraries, archivists, and book historians think about preservation.

The significance of this story extends well beyond Anthropic's specific practices. It captures a broader tension in the generative AI industry between the insatiable demand for high-quality training data and the human and cultural infrastructure that data comes from. Anthropic's $1.5 billion settlement with authors in September 2025—the largest publicly disclosed copyright recovery in U.S. history—already signaled how contentious the sourcing of training material has become. But the detail that Anthropic went so far as to buy and destroy physical books, rather than simply license digital editions or work with libraries on non-destructive digitization, illustrates the extent to which AI labs have prioritized speed and control over data acquisition, even when less destructive alternatives (such as the kind of careful digitization long practiced by institutions like the Internet Archive or academic libraries) exist.

More broadly, the episode reflects how the AI industry's data hunger increasingly intersects with cultural heritage and preservation communities that have historically operated under very different norms and incentives. Libraries and archives have spent decades developing practices designed to preserve originals while creating access copies; Anthropic's approach, by contrast, treated preservation as incidental to the goal of building a defensible legal claim to "owned" training data. James's essay implicitly asks what is lost—materially, historically, and culturally—when books are treated as disposable inputs to a computational pipeline. As AI companies continue to race for proprietary, legally clean training corpora, the fate of the physical books consumed along the way is likely to remain a flashpoint, especially as authors, librarians, and preservationists push back against a model of "progress" that quite literally destroys the object it claims to be learning from.

Read original article →