Detailed Analysis
Anthropic's approach to acquiring training data for Claude involved a large-scale physical book digitization effort that came to light through court filings in the copyright lawsuit brought by a group of authors against the company. According to testimony and evidence disclosed during litigation, Anthropic purchased millions of used print books, then physically destroyed them—cutting off their spines and scanning the loose pages—in order to convert them into machine-readable text for training its large language models. The company reportedly hired former Google Books executive Tom Turvey to lead this effort, and internal records indicate Anthropic viewed this "destructive scanning" process as a way to create a defensible legal position, since it involved books the company had legitimately purchased rather than pirated copies.
This detail matters significantly because it sits at the center of one of the most consequential legal battles shaping the future of AI copyright law. In a landmark ruling in the case Bartz v. Anthropic, U.S. District Judge William Alsup drew a sharp distinction between the legality of training on lawfully acquired, scanned books versus training on pirated digital copies obtained from shadow libraries like Library Genesis (LibGen) and Books3. Alsup found that Anthropic's use of purchased and destructively scanned books for AI training likely qualified as fair use, reasoning that the transformation of physical books into a research and training corpus was sufficiently transformative. However, he ruled that Anthropic's earlier practice of downloading millions of pirated books to build a general-purpose library constituted copyright infringement, exposing the company to potentially massive statutory damages—ultimately resulting in a $1.5 billion settlement, one of the largest copyright settlements in history.
The book-shredding revelation illustrates the lengths to which AI companies have gone to secure high-quality training data while attempting to navigate murky and rapidly evolving copyright law. High-quality, long-form text from published books is considered extremely valuable for training language models because it offers coherent, well-edited, and diverse writing compared to much of the scraped internet content used in earlier training pipelines. By physically buying and destroying books rather than pirating digital scans, Anthropic sought to align its data acquisition with legal precedents from cases like Authors Guild v. Google Books, which established that scanning physically-owned books for search and analysis purposes could constitute fair use under certain conditions.
This episode reflects a broader tension defining the current era of AI development: the voracious data requirements of large language models colliding with intellectual property rights and the economic interests of creators. As foundation model companies including OpenAI, Meta, and Google face their own lawsuits over training data sourcing, Anthropic's book-scanning saga—and the resulting settlement—is likely to serve as an influential precedent, pushing the industry toward more expensive, licensed, or physically-sourced data acquisition strategies rather than reliance on freely scraped or pirated content. It also underscores how courts are beginning to differentiate between the legality of the training process itself and the legality of how the underlying data was obtained, a distinction that will likely shape licensing negotiations, damages calculations, and compliance strategies across the AI industry for years to come.
Read original article →