← Reddit

Anthropic is destroying old rare books without digitally preserving them

Reddit · Used-Nectarine5541 · July 30, 2026
A commenter expressed strong opposition to Anthropic's alleged destruction of rare books without digital preservation, characterizing the action as morally corrupt. The commenter emphasized the critical importance of preserving knowledge and questioned the rationale for such practices.

Detailed Analysis

A Reddit post alleging that Anthropic is "destroying old rare books without digitally preserving them" has circulated with strong emotional language but virtually no verifiable detail, sourcing, or corroborating evidence. The post itself reads as a reaction of alarm rather than a documented report—it contains no citation, no link to primary reporting, no quotes from Anthropic, and no description of which books, how many, or under what circumstances any destruction allegedly occurred. Without additional context confirming the claim, it's important to treat this as an unverified allegation rather than an established fact.

What is publicly known is that Anthropic has previously been involved in a large-scale book-scanning effort tied to its AI training pipeline. In legal proceedings (notably the Bartz v. Anthropic copyright lawsuit), it emerged that the company purchased millions of physical books, cut their bindings, and scanned the pages to convert them into digital text for training its Claude models—a process that is inherently destructive to the physical book but explicitly done to create a digital copy. That scanning-and-destroying method was actually cited by Anthropic's own legal team as evidence of good-faith "fair use," since the company was said to be destroying one copy (the physical book) while creating exactly one replacement copy (the digital scan), rather than mass-producing pirated copies for distribution. A federal judge's June 2025 ruling on the case found this practice of scanning purchased print books to be transformative and likely protected by fair use, distinguishing it from Anthropic's separate and more legally fraught practice of downloading pirated ebooks from shadow libraries, which the same judge said was not protected.

The confusion in the Reddit post likely stems from a garbled or exaggerated understanding of this scanning process. The physical destruction of a book during scanning is standard practice in many large-scale digitization projects (Google Books used similar destructive scanning methods in some cases), and the entire point of Anthropic's process—as described in court filings—is to preserve the content digitally by converting it into machine-readable text, not to destroy knowledge without preserving it. If the claim in the post refers to something else entirely—perhaps a specific incident involving genuinely rare or irreplaceable volumes destroyed without any digital backup—no such incident has been substantiated in available reporting, and the burden would be on someone making that claim to provide documentation, since it would represent a significant escalation beyond what has been publicly reported about Anthropic's training data practices.

This episode is illustrative of a broader pattern in AI discourse: emotionally charged claims about AI companies' data practices spread quickly on social platforms like Reddit, often outpacing verification. It also reflects genuine and legitimate public anxiety about how AI companies acquire training data—an anxiety rooted in real controversies, including the pirated-ebook allegations in the Bartz case, disputes over copyrighted material scraped from the web, and broader tensions between authors, publishers, and AI developers over compensation and consent. Anthropic has faced real legal and reputational consequences tied to its data acquisition methods, including a $1.5 billion settlement with authors over the pirated books used in training. So while this specific claim about destroying rare books without preservation appears unsubstantiated as stated, it exists within a larger, legitimate conversation about transparency, consent, and the environmental and cultural costs of how large language models are built.

Read original article →