The Bartz/Anthropic Settlement is still a hot mess but at least it is the largest known copyright recovery of all time (for piracy)?

57 views
Skip to first unread message

Georgia Jenkins

unread,
Aug 11, 2026, 7:12:37 AM (5 days ago) Aug 11
to ipkat_...@googlegroups.com
Home / AI / Anthropic / Artificial Intelligence / class action / Copyright infringement / Georgia Jenkins / Illegal downloads / napster / piracy / settlement / USA The Bartz/Anthropic Settlement is still a hot mess but at least it is the largest known copyright recovery of all time (for piracy)?

The Bartz/Anthropic Settlement is still a hot mess but at least it is the largest known copyright recovery of all time (for piracy)?

Over a year ago, this Kat reported on Anthropic’s motion for summary judgment that cited fair use (IPKat here). In August 2024, Andrea Bartz, Charles Graeber and Kirk Wallace Johnson alleged that Anthropic’s unauthorised use of ‘pirated’ copies and purchased print books (later digitised) infringed copyright in their works. While (the now retired) Judge Alsup found that the latter benefited from fair use, the remaining ‘pirated copies’ required a trial.

Ahead of the trial, not only did the lawsuit become a certified class action, but the parties reached a settlement. Anthropic agreed to pay at least US$1.5 billion plus interest to authors for 482,460 books (roughly US$3,000 per work). Additionally, in exchange for destroying the pirated datasets, Anthropic avoided litigation for related conduct up to 25 August 2025. Following a long list of questions relating to settlement administration (here and here), Judge Alsup preliminarily approved the settlement in September 2025.

While this Kat questioned the worth of a book used to train LLM models, others declared that this was ‘the largest copyright recovery of all time’. The Anthropic Copyright Settlement Website provided a tool for rightsholders to check whether their work was listed and submit a claim by the end of March 2026. The deadline for opting out of the settlement and objection was February 2026.

The final settlement approval hearing occurred in May 2026 where it was submitted that 440,490 of the 482,460 were works eligible for the fund. Judge Martínez-Olguín approved the settlement and the order was published at the end of July.
This Kat recalls copyright once referred to as a 'dumpster fire'.
Hot mess feels fitting for the Settlement Fund.
Photo by Olivia Brewer on Unsplash

Books, lists of works and a certified class action
Before Judge Alsup preliminarily approved the settlement, there was a motion to certify two classes of copyright owners: “Pirated Books Class”, and a “Scanned Books Class”. It was submitted by the plaintiffs that the “Pirated Books Class”, the subject of the settlement, comprised two potential subsets across three pirate libraries: (1) Books3; and (2) Library Genesis (LibGen) and Pirate Library Mirror (PiLiMi). All represent a spectrum of unauthorised copying: downloading Books3; copying LibGen via BitTorrent; and torrenting a new version of LibGen that third parties created to mirror and improve PiLiMi. Judge Alsup summed it up as ‘Napster-like’ copying.

Earlier on, Judge Alsup had denied the certification for a ‘Books3 Pirated Books Class’ as it lacked sufficient metadata and comprised fewer complete files of content. This would have made the identification of titles and authors too problematic. Ahead of the final hearing, the plaintiffs, again, submitted that the certified class comprised the Books3 subset. Referring to Judge Alsup’s preliminary approval, which cited the certification order, Judge Martínez-Olguín affirmed the exclusion of Books3 from the list of works. In comparison, Nazemian v. NVIDIA, in which NVIDIA allegedly trained its LLMs using ‘the Pile’ that includes Books3, will be one to watch as it starts discovery.

The insufficient nature of Books3 also links to Anthropic’s use of LibGen, PiLiMi, and ultimately, the scanned books. Anthropic had overlapping copies of some books which impacted LLM memorisation and performance. Alongside concerns relating to ‘the breadth and depth of categories selected’, the likelihood of regurgitating copyright- protected training content, and the impact of removing ‘pirate-sourced copies’, Anthropic sought to identify books with ‘old-books quality’. This process was later used by the court to define which books form part of the list of works. As Anthropic focused on purchasing books with an International Standard Book Number (ISBN) or Amazon Standard Identification Number (ASIN), such ‘books’ were required to be registered with the United States Copyright Office within five years of the work’s first publication and before being downloaded by Anthropic, or within three months of first publication.

A ‘fair, reasonable and adequate’ proposed settlement
In the US, class action settlements require the court’s approval pursuant to Federal Rule of Civil Procedure 23(e). Following the approval of notice of the proposal to all class members, the court evaluated whether it was ‘fair, reasonable, and adequate’. This requires considering whether: the class representatives and class counsel have adequately represented the class; the proposal was negotiated at arm’s length; the relief provided for the class is adequate; and the proposal treats class members equitably relative to each other.

Judge Martínez-Olguín emphasised the reasonable nature of the settlement and highlighted the ‘complex, expensive, lengthy and risky’ nature of future litigation. Additionally, the approximate per-work payment of US$3,000 was ‘four times the minimum statutory damages for wilful infringement’. Referring to the Hanlon factors that evaluate the fairness of a class action settlement, Judge Martínez-Olguín listed the strength of the plaintiff’s case, the risk of maintaining class action status throughout the trial, the extent of discovery completed and the stage of proceedings, the experience and views of counsel, and the reaction of the class members as indicative of its fairness.

The order explained that the settlement provides ‘substantial benefits […] in light of the novel claims asserted’. While there was likely a strong case on downloading, Judge Alsup’s earlier order weakened arguments related to training LLMs as infringement. Lastly, the Class response was ‘overwhelmingly favourable’ amounting to a claim rate of at least 91.3% of works. Most interesting for this Kat, were the objections based on non-monetary relief, including novel licensing schemes, source attribution in Anthropic’s outputs, deletion of AI models, and abolishing the use of scanned books for training. They were overruled as Judge Alsup narrowed the copyright infringement claims to Anthropic downloading ‘pirate-sourced’ copies, not using copies to train LLMs.

Plan of distribution, allocation and claims process
The court also approved the Parties’ proposal for allocating the Net Settlement Fund. Here, valid claimants will receive a pro rata per-work share, to be divided among copyright owners on either elected default splits or a percentage split determined by contracts of publishing agreements. If any funds remain after all valid claims are paid, they will be redistributed among the settlement class members until it is economically unfeasible to do so. Subject to court approval, any final balance will be distributed mirroring the above approach.

However, there were objections relating to the plan allocation (here and here) alluding to power imbalances within book publishing alongside what some have described as ‘lawyers [..] rushing [a] historic settlement to seize US$320 million in fees’. One author wrote in their objection that, ‘every dollar that Counsel takes from the Settlement fund is one that is not given to those actually harmed’. While initially counsel requested 20% of the settlement in fees, Judge Martínez-Olguín reduced it to 6.8%, amounting to $101,561,111. To address some of these concerns, the order also includes a ‘post-distribution accounting’ mechanism.

Comment
This story is far from over. Not only does the settlement exclude the certification Books and claims relating to AI outputs and future works, but there were 350 valid opt-outs over 1,802 works. Still, Judge Alsup’s finding that training LLMs is transformative casts a large shadow over future AI copyright infringement proceedings with the result that, at least in the US, they must use Napster-like rhetoric. Perhaps more remarkably, the only reason that ‘pirate-sourced copies’ came into view was because, to benefit from fair use, Anthropic outlined the process of training LLMs which includes where they sourced those copies.

Without Anthropic’s disclosures, proving the ‘pirate-sourced’ nature of the copies may have cost more than US$101,561,111. At this point, such costs frustrate the very purpose of a class action settlement fund: collective redress for the exploitation of creativity when training LLMs. Again, this Kat suspects that copyright law may be poorly suited to address the cumulative harm that humanity continues to face. A class action lawsuit seemingly mirrors this type of harm. However its sky high legal fees and restriction to pirated copies, makes its response patchy, at best. The settlement evidences that the copyright system has 'been weighed, measured, and found wanting' (Count Adhemar in this Kat’s favourite Heath Ledger film, ‘A Knight’s Tale’).
Do you want to reuse the IPKat content? Please refer to our 'Policies' section. If you have any queries or requests for permission, please get in touch with the IPKat team.
Reply all
Reply to author
Forward
0 new messages