Detailed Analysis
Anthropic's practice of purchasing used physical books in bulk, cutting off their bindings, scanning each page, and then discarding the paper originals has emerged as a central and somewhat unusual element of its legal strategy in the ongoing copyright litigation brought by authors over training data used to build Claude. Rather than downloading pirated digital copies from shadow libraries—the practice that got Anthropic into serious legal trouble and exposed it to potential liability in the hundreds of millions or billions of dollars—the company has apparently sought to establish a cleaner paper trail by acquiring physical books through legitimate retail channels, destructively scanning them into digital form, and treating that process as the functional equivalent of format-shifting content it legally owns. This approach featured prominently in Judge William Alsup's ruling in the Bartz v. Anthropic case, where he distinguished between Anthropic's use of pirated copies (found to weigh against fair use) and its practice of buying and scanning physical books (which he treated far more favorably, comparing it to a library patron converting formats for personal use).
The strategy matters because it sits at the heart of how courts are beginning to parse fair use doctrine as applied to AI training. Alsup's June 2025 decision was among the first substantial judicial rulings to address whether training large language models on copyrighted text constitutes transformative fair use, and he split the question into two distinct threads: the training use itself, which he deemed transformative and lawful, and the acquisition method, which mattered enormously for liability. Buying a physical book and destructively scanning it, in his reasoning, resembles legitimate personal format-shifting akin to ripping a CD to MP3, whereas downloading the same book from a pirate site to build a permanent digital library is a separate, unauthorized act of reproduction regardless of what happens afterward. This distinction has left Anthropic facing a massive settlement—reportedly around $1.5 billion—tied specifically to the pirated-book portion of its training corpus, while the scan-and-destroy method has effectively become a template other AI developers are said to be emulating to insulate themselves from similar claims.
The physical mechanics of the process—slicing bindings off secondhand books, feeding loose pages through high-speed scanners, and discarding the resulting paper waste—have drawn public attention and some discomfort, both for the literal destruction of physical texts and for what it reveals about the scale of data acquisition underpinning modern AI systems. Reports indicate warehouses processing books by the truckload, reflecting how enormous the appetite for training data has become as models grow larger and companies race to secure legally defensible sources of text. The imagery of destroyed books also carries symbolic weight in a debate that pits authors, publishers, and the Authors Guild against trillion-dollar technology companies, feeding into broader anxieties about AI's relationship to human creative labor and cultural heritage.
More broadly, this episode reflects how AI companies are scrambling to retrofit legally sound provenance onto training pipelines that were often built quickly and with limited attention to intellectual property rights during the early stages of the LLM race. As courts, regulators, and litigants sharpen their scrutiny of training data sourcing, companies like Anthropic, OpenAI, Meta, and Google are all navigating similar exposure, and the outcomes of cases like Bartz v. Anthropic are likely to set precedents shaping how the entire industry acquires data going forward. The scan-and-destroy approach, however legally expedient, underscores a deeper tension in the AI industry: the drive to build ever-larger models on vast troves of human-created text is colliding with copyright frameworks never designed for machine-scale consumption, forcing companies into legally cautious but practically strange workarounds—destroying physical books to seemingly launder their content into a permissible digital form.
Read original article →