Anthropic paid $1.5 bn to train Claude on illegal literature

Tue Jul 28 2026
Rajesh Sharma (2323 articles)
Anthropic paid $1.5 bn to train Claude on illegal literature

A US federal judge has granted final approval to Anthropic’s significant $1.5 billion settlement with authors who alleged that the artificial intelligence company utilised pirated books for training its AI assistant and large language model, Claude. The settlement, sanctioned by a judge in San Francisco, arises as AI developers such as Anthropic, OpenAI, and Meta encounter an increasing array of lawsuits from authors, publishers, news organisations, and other copyright holders regarding the utilisation of copyrighted material for training large language models. Anthropic reported that over 91 percent of eligible authors and publishers have successfully claimed their portion of the payout. “We reached this settlement in 2025, after the court’s landmark ruling that training AI on books is fair use under copyright law, which remains the law today,” Anthropic deputy general counsel Aparna Sridhar said in a statement. “We are pleased that more than 91 percent of authors and publishers covered by the settlement have claimed their share of the payment, and we’re looking forward to bringing this matter to a close.”

The lawsuit was initiated in 2024 by authors Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson, who claimed that Anthropic appropriated millions of books without authorisation during the development of Claude. Court filings indicated that the company established an extensive digital library utilising two distinct sources. It acquired millions of physical books that were subsequently digitised. In a separate instance, it acquired over seven million books from online piracy repositories, such as Library Genesis and Pirate Library Mirror. US District Judge William Alsup made a definitive differentiation between the books that Anthropic had acquired through legal means and those it had procured via piracy. He ruled that training AI models using legally acquired books constituted fair use under US copyright law, as the court determined that the books were not being replicated to supplant the originals, but rather were utilised to instruct the AI model in the mechanics of language. That rendered the use “exceedingly transformative” under US copyright law. He also maintained that digitising legally purchased print books into digital format for internal AI training was permissible, considering it a format conversion rather than unlawful copying. However, concerning pirated books, he ruled that downloading and maintaining a permanent internal library constructed from illegally obtained books constituted copyright infringement, even if those books were subsequently utilised for AI training. The ruling positioned Anthropic at risk of incurring substantial statutory damages should the case proceed to trial. Instead, the company has consented to establish a settlement fund totalling a minimum of $1.5 billion, with qualifying authors and publishers anticipated to receive approximately $3,000 for each eligible work prior to the deduction of legal fees and other expenses.

While the settlement focused on pirated books, the litigation also uncovered the methods by which Anthropic acquired legally purchased physical books for the purpose of AI training. Recently unsealed court documents revealed that, as part of an internal initiative referred to as Project Panama, the company allocated tens of millions of dollars to acquire used books from booksellers and libraries. Following this acquisition, the company proceeded to remove the bindings using industrial cutting equipment, scanned each page with high-speed production scanners, and subsequently recycled or disposed of the physical copies. Internal planning documents, initially disclosed by The Washington Post, characterised the initiative as an endeavour to “destructively scan all the books in the world” and directed employees to refrain from public discussions regarding the project. Within a year, millions of books were both acquired and disposed of. Judge Alsup concluded that it was lawful because Anthropic had legally purchased the books before converting them into digital copies for internal use.

According to a report, an independent technology publication that focuses on digital rights and internet culture, companies providing books to AI developers have been increasingly acquiring physical books in bulk, including rare and out-of-print editions, for the purpose of destructive scanning. Among the companies operating in this sector is ISBNdb, a commercial book database that has traditionally supplied bibliographic information to publishers, booksellers, and libraries. However, it has progressively redefined its role as a provider of physical books for AI training. According to the report, ISBNdb contends that books published prior to the extensive use of generative AI offer cleaner training data, as they lack AI-generated content. The company enables the acquisition of up to one million books for AI training purposes and recognises that the mass destruction of books may provoke significant public opposition. Booksellers in Europe and North America have reported significant purchases of specialised academic titles, foreign-language books, and rare editions. This trend raises concerns that valuable physical copies may permanently vanish from circulation as AI companies compete to secure training material.

The settlement concludes Anthropic’s contention regarding these pirated books; however, it leaves unresolved the broader discourse surrounding the permissibility of AI firms utilising copyrighted content for model training. Some authors and publishers have chosen to withdraw from the settlement and are actively pursuing independent lawsuits against Anthropic. The agreement is limited to past claims concerning books that are part of the class list and does not inhibit future legal actions regarding new copyright matters or purported infringements stemming from Claude’s outputs. The ruling also offers an initial framework for how courts might interpret AI training in relation to copyright issues. It indicates that judges exhibit a greater readiness to endorse AI models trained on legally acquired material compared to those utilising datasets derived from pirated books. As litigation progresses against Anthropic, OpenAI, Meta, and other AI developers, this case is anticipated to establish a significant early precedent in delineating the permissible use of copyrighted materials for training artificial intelligence systems.

Rajesh Sharma

Rajesh Sharma

Rajesh Sharma is Correspondent for Stock Market of South East Asia based in Mumbai. He has been covering Asian markets for more than 5 years.