Detailed Analysis
Anthropic's book-buying-and-destruction operation stems directly from the copyright infringement lawsuit brought by a group of authors, which the company settled in 2025 for $1.5 billion—one of the largest copyright settlements in history. That litigation centered on Anthropic's use of pirated digital book repositories (like Books3 and LibGen) to train its Claude models. As part of its legal and strategic response, Anthropic pivoted to acquiring physical print books in bulk, destructively scanning them into digital text, and then discarding the physical copies. The internal memo language cited in the article—"we do not want it to be known that we are pursuing this project"—reveals a company acutely aware that this practice, while arguably more legally defensible than piracy, still carries reputational risk if publicized.
The legal logic behind this approach is notable: U.S. copyright law and the fair use doctrine have historically treated the act of purchasing a physical book and converting it to another format for internal use more favorably than mass-scale piracy, particularly when the training use is transformative. Judge William Alsup's ruling in the Anthropic case last year drew exactly this distinction, finding that training an AI on lawfully acquired books could qualify as fair use, while training on pirated copies could not. This created a strong incentive for Anthropic and peer companies to establish a paper trail of legitimate acquisition—buying physical books, often rare or out-of-print titles not available in convenient digital form—even though the process of shredding or destroying books after scanning strikes many observers as wasteful, ethically fraught, and symbolically troubling given the cultural and historical value of physical texts.
This story fits into a broader pattern of AI labs scrambling to secure vast, legally defensible training corpora in the wake of successful and pending litigation from authors, publishers, news organizations, and artists. Companies including OpenAI, Meta, and Google have faced similar suits over the sourcing of their training data, and all have had to reckon with the tension between insatiable demand for high-quality text data and the legal/ethical constraints on how that data can be obtained. Anthropic's approach—paying for books rather than pirating them—represents an attempt to thread this needle, but the secrecy described in the memo, combined with the physically destructive nature of the scanning process, undercuts the company's public positioning around responsible and ethical AI development.
More broadly, the episode illustrates how the economics of large language model training have created strange, almost analog bottlenecks in an otherwise digital industry: books that exist only in physical form, especially rare, out-of-print, or niche titles, have become valuable raw material, spurring AI companies to essentially strip-mine libraries and used bookstores. It also underscores growing scrutiny of AI companies' data practices generally, as journalists, authors' guilds, and regulators increasingly demand transparency about what content trains these systems and under what terms. Anthropic's stated desire for secrecy, even while operating within a legally sanctioned process, is likely to fuel further public and regulatory skepticism about the industry's data-sourcing practices, adding pressure for clearer disclosure norms as AI training data litigation continues to unfold across the sector.
Read original article →