← Reddit

Destruction of books? Project Panama.

Reddit · ZealousidealWafer383 · July 29, 2026
A forum post alleges that Anthropic destroyed millions of books to scan them for AI model training and questions why the well-funded company did not employ non-destructive scanning methods. The post raises concerns about whether rare books and religious texts were among those destroyed and requests public clarification from Anthropic on the matter.

Detailed Analysis

The Reddit post raises questions about a practice that surfaced during Anthropic's copyright litigation: the company's purchase and physical destruction of millions of print books in order to digitize them for use as training data, an initiative reportedly referred to internally as "Project Panama." Rather than licensing digital editions or negotiating with publishers, Anthropic acquired used physical copies in bulk, cut off their bindings, scanned each page, and discarded the paper remnants—retaining only the digital text for use in training its Claude models. The practice came to light through court filings in the Bartz v. Anthropic case, a lawsuit brought by authors alleging that Anthropic infringed their copyrights by training on pirated and scanned books without permission or compensation.

The destructive scanning method is not, as the original poster speculates, a matter of technological necessity—high-speed destructive scanners are simply cheaper and faster than non-destructive alternatives, and buying secondhand books in bulk sidesteps the friction of licensing negotiations with publishers and authors. Legally, this approach became central to Anthropic's defense: the company argued that because it had purchased legitimate physical copies before converting them to digital form, its use of that scanned content for training fell under fair use, similar to a person buying a book, extracting insights from it, and creating something new. In its June 2025 ruling, the federal judge overseeing the case largely agreed, finding that training an LLM on legally purchased books constituted transformative fair use—while also ruling that Anthropic's separate practice of downloading pirated books from shadow libraries did not enjoy the same protection. That distinction proved significant: Anthropic ultimately agreed to a $1.5 billion settlement, one of the largest copyright settlements in history, largely tied to the pirated-book claims rather than the destructive-scanning practice itself.

The controversy illuminates a broader tension in the AI industry: the voracious data appetite of large language models colliding with authors' and publishers' rights, and the ad hoc, sometimes extralegal methods companies have used to satisfy that appetite. Destroying physical books to create private digital corpora sits in an ethically ambiguous middle ground—distinct from outright piracy, since the books were purchased, but still troubling to many because the resulting digital assets are never shared, resold, or returned to public circulation; they simply vanish into a proprietary training pipeline. This echoes long-standing debates around Google Books, HathiTrust, and other mass-digitization projects, where the difference between "scanning for search/reference" and "scanning to train a commercial product that competes with the source material" has been legally and morally contested.

For everyday users and observers, this episode is part of a larger reckoning happening across the generative AI sector in 2025–2026, as courts, regulators, and the public increasingly scrutinize how foundation models are built. Anthropic, despite branding itself as a safety-focused and relatively transparent AI lab compared to peers like OpenAI or Meta, has faced real consequences for its data-sourcing choices, showing that even companies with strong public safety narratives are not exempt from copyright accountability. The size of the settlement and the visibility of practices like Project Panama are likely to influence how future AI labs approach training data acquisition—pushing the industry toward more formal licensing arrangements with publishers and rights holders rather than mass ad hoc digitization, destructive or otherwise.

Read original article →