← Google News

Harry Potter publisher to receive millions in Anthropic copyright settlement - The Guardian

Google News · July 22, 2026
Harry Potter publisher to receive millions in Anthropic copyright settlement The Guardian [truncated: Google News RSS provides only a snippet, not full article

Detailed Analysis

Anthropic's copyright settlement over its use of pirated books to train Claude continues to generate significant financial consequences, with reports indicating that Bloomsbury—the publisher behind the Harry Potter series in the UK and numerous other bestselling titles—stands to receive millions of dollars as part of the compensation process. This development follows the landmark $1.5 billion settlement Anthropic reached in 2025 with a class of authors and publishers who alleged the company had downloaded and used pirated copies of their works, sourced from shadow libraries like Books3 and LibGen, to train its large language models. That settlement, approved by U.S. District Judge William Alsup in the Northern District of California, is widely regarded as the largest copyright recovery in publishing history, working out to roughly $3,000 per infringed work across an estimated 500,000 titles.

The case originated from a lawsuit brought by authors including Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson, who argued that Anthropic's practice of acquiring and retaining pirated text datasets constituted willful copyright infringement, regardless of whether the resulting AI outputs were transformative. Judge Alsup's earlier rulings drew an important distinction: he found that training an AI model on legally acquired copyrighted books could plausibly qualify as fair use because it involved a transformative process akin to how a human learns from reading, but he was far less sympathetic to Anthropic's acquisition of millions of books through piracy, calling that practice indefensible regardless of downstream use. This distinction has become an influential legal template as courts grapple with how copyright law applies to AI training data.

The involvement of major publishers like Bloomsbury—rather than just individual authors—signals the scale and reach of the settlement's claims process, as rights holders across the publishing industry, including those representing globally recognized franchises, seek compensation for unauthorized use of their catalogs. Bloomsbury's stake in the settlement underscores how even publishers of enormously commercially successful properties were not exempt from having their works swept into the datasets used to train generative AI systems, raising questions about how thoroughly AI companies vetted their training data sourcing during the industry's rapid buildout phase.

This settlement carries broader implications for the AI industry, which has faced a wave of copyright litigation from authors, news organizations, visual artists, and musicians alleging similarly unauthorized use of copyrighted material to train generative models. Anthropic's payout, alongside ongoing suits against OpenAI, Meta, and Stability AI, signals that courts and rights holders are increasingly willing to impose substantial financial penalties on AI developers for how they source training data, even as the legal question of whether AI training itself constitutes fair use remains contested and unsettled. For an industry that has largely operated on the assumption that scraping vast quantities of text and images was a low-risk cost of doing business, the size of the Anthropic settlement may force competitors to reassess their own data provenance practices, potentially reshaping how future AI systems are trained and licensed, and increasing pressure toward negotiated licensing deals with publishers rather than unilateral data scraping.

Read original article →