← Google News

Judge approves a $1.5B Anthropic settlement over books used to train Claude - KMTV 3 News Now

Google News · July 23, 2026
Judge approves a $1.5B Anthropic settlement over books used to train Claude KMTV 3 News Now [truncated: Google News RSS provides only a snippet, not full article

Detailed Analysis

A federal judge has granted final approval to a $1.5 billion settlement between Anthropic and a group of authors who alleged the company illegally used pirated copies of their books to train its Claude AI models. The case, which originated in a class-action lawsuit filed by authors including Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson, centered on Anthropic's use of shadow libraries such as Library Genesis (LibGen) and Pirated Books Library (PiLiMi) to build a dataset for training its large language models. The settlement, believed to be the largest publicly reported copyright recovery in U.S. history, will compensate authors and rights holders whose works were included in these pirated datasets, with payments reportedly amounting to roughly $3,000 per book across an estimated 500,000 or more titles.

The legal proceedings drew a notable distinction between two separate practices: purchasing and scanning books versus downloading pirated copies from illegal repositories. In an earlier ruling, U.S. District Judge William Alsup found that Anthropic's practice of training on legally purchased books could qualify as fair use under copyright law, a determination that offered some reassurance to AI developers regarding the use of lawfully acquired training material. However, the judge was far less sympathetic to Anthropic's admitted use of pirated texts obtained through LibGen and similar sites, ruling that this conduct exposed the company to significant liability regardless of how the material was ultimately used. This bifurcated outcome—protecting fair use for legitimately sourced content while punishing piracy—has emerged as an important early precedent shaping how courts may handle the wave of AI copyright litigation now working through the federal court system.

The financial scale of the settlement underscores the escalating legal and financial risks facing AI companies as they confront the consequences of how their foundation models were trained. Anthropic, backed by billions in investment from Amazon, Google, and other major technology players, built Claude into one of the leading competitors to OpenAI's ChatGPT and Google's Gemini, but the settlement now stands as a costly reminder that the company's early data-sourcing decisions carry lasting liability. For an industry that has largely operated under a "move fast and ask forgiveness later" approach to data acquisition, this case signals that courts are willing to impose severe financial penalties when companies can be shown to have knowingly relied on pirated content, even if other training practices are deemed permissible.

Beyond Anthropic, the settlement carries broad implications for the entire generative AI sector, which faces a growing docket of similar lawsuits from authors, publishers, visual artists, and music labels against companies including OpenAI, Meta, Microsoft, and Stability AI. The ruling and settlement amount are likely to influence litigation strategy industry-wide, potentially pushing AI developers toward more rigorous data licensing agreements and away from reliance on unauthorized or pirated repositories. It also strengthens the negotiating position of publishers and authors' guilds, who have long argued that AI training constitutes a form of mass copyright infringement deserving compensation. As regulators and courts continue grappling with how existing copyright law applies to machine learning, this settlement offers one of the clearest signals yet that the era of unchecked data scraping for AI training is drawing to a close, with real financial consequences now attached to how companies source the material that powers their models.

Read original article →