← Google News

How Anthropic ended up paying $1.5 bn for training Claude on pirated books - Business Standard

Google News · July 28, 2026
How Anthropic ended up paying $1.5 bn for training Claude on pirated books Business Standard [truncated: Google News RSS provides only a snippet, not full article

Detailed Analysis

Anthropic's agreement to pay $1.5 billion to settle a class-action lawsuit brought by authors and publishers marks one of the largest copyright settlements in the history of the technology industry, and it stands as a landmark moment in the ongoing legal reckoning over how AI companies acquired the data used to train their large language models. The case centered on allegations that Anthropic downloaded millions of pirated books from shadow libraries such as Library Genesis (LibGen) and Pirate Library Mirror (Books3-adjacent sources) to build the training corpus for Claude, rather than licensing the material or acquiring it through legitimate channels. Judge William Alsup of the Northern District of California had already ruled earlier in 2025 that while training an AI model on copyrighted books could qualify as fair use, the act of pirating those books in the first place was not protected and exposed Anthropic to liability for illegal acquisition, distinct from the training process itself.

The settlement's scale—reportedly covering around 500,000 works at roughly $3,000 per title—reflects the sheer volume of copyrighted material implicated and sets a financial benchmark that other AI companies facing similar litigation will now have to reckon with. This nuance is critical: Anthropic did not simply lose a broad fight over whether AI training itself infringes copyright. Instead, it was the sourcing method, obtaining books from piracy sites rather than purchasing or licensing them, that proved legally indefensible. This distinction gives some legal breathing room to AI developers who can demonstrate legitimate acquisition of training data, while simultaneously raising the stakes for those who cannot.

The case matters well beyond Anthropic because nearly every major AI lab, including OpenAI, Meta, and Google, has faced or is currently facing comparable lawsuits from authors, artists, musicians, and publishers alleging that copyrighted works were scraped or pirated to train generative AI systems. Anthropic's settlement offers a template and a cautionary benchmark: it demonstrates that even a company that has cultivated a reputation for prioritizing AI safety and responsible development is not insulated from liability tied to how its foundational data was sourced. For rights holders, the outcome validates years of advocacy arguing that "fair use" cannot serve as a blanket defense when the underlying materials were obtained illegally, and it likely emboldens further litigation and demands for licensing deals across the industry.

More broadly, this settlement accelerates a shift already underway in the AI industry toward formal licensing arrangements with publishers, news organizations, music labels, and other content creators, as companies seek to avoid the legal and reputational risks exposed by cases like this one. It also underscores the tension between the enormous data appetite required to train competitive frontier models and the intellectual property rights of the creators whose work makes that training possible. As regulators and courts continue to grapple with how copyright law applies to AI, Anthropic's $1.5 billion payout signals that the era of using pirated or unlicensed content with impunity is closing, pushing the industry toward a more costly, but more legally defensible, data-sourcing model going forward.

Read original article →