Lawsuit tracker

N.D. Cal. · Settled

Bartz v. Anthropic

Anthropic just closed the largest copyright settlement in U.S. history — $1.5 billion for training Claude on pirated books. The case didn't decide whether AI training is fair use. It decided how you got the data.

Key facts

  1. Anthropic downloaded roughly seven million books from the pirate libraries LibGen and PiLiMi.
  2. Judge William Alsup ruled in June 2025 that training on lawfully acquired, purchased-and-scanned books was fair use, but that downloading and retaining pirated copies was not.
  3. The certified class covered rightsholders in about 482,460 works, with theoretical statutory exposure above $70 billion.
  4. The parties settled for $1.5 billion — about $3,000 per covered work — the largest copyright settlement on record.
  5. Judge Araceli Martínez-Olguín granted final approval on July 20, 2026 and cut the requested $187.5 million attorneys’ fee award to about $101.6 million.
  6. The settlement releases past conduct only. It sets no binding precedent and includes no forward license or output coverage.
PartiesAndrea Bartz, Charles Graeber and Kirk Wallace Johnson, on behalf of a certified class of rightsholders in roughly 482,460 books v. Anthropic PBC
CourtU.S. District Court for the Northern District of California
Docket24-cv-05417
Filed2024-08-19
Content typeBooks — trade and academic titles from the LibGen and PiLiMi pirate libraries, plus purchased-and-scanned print copies
StatusAs of July 2026, settled. Judge Araceli Martínez-Olguín granted final approval of the $1.5 billion class settlement on July 20, 2026 — roughly $3,000 per book, with the court noting the deal sets no binding precedent.

What was alleged

Andrea Bartz, Charles Graeber and Kirk Wallace Johnson sued Anthropic PBC in the Northern District of California on August 19, 2024. Their claim was narrow: direct copyright infringement. Anthropic had downloaded their books from pirate libraries and copied them to train the Claude models.

The suit named three authors. It became a certified class covering rightsholders in roughly 482,460 books — one of the largest copyright classes ever certified against an AI developer.

Two libraries, seven million books

Discovery in the case surfaced the mechanics behind Claude’s training corpus. Anthropic had downloaded about seven million books from two pirate libraries, LibGen and PiLiMi. Separately, and on a different legal footing, it had bought millions of print books and scanned them in-house.

That second pile of books — purchased, then digitized — is the detail that ended up mattering most. It gave Judge William Alsup two different fact patterns to rule on inside a single case.

The split that mattered

In June 2025, Alsup split the case in two. Training a model on lawfully acquired, purchased books was fair use, full stop. But downloading and retaining pirated copies — even books Anthropic never used for training, even books it later bought legitimately — was not fair use. The pirated copies were the injury, not the training run.

Alsup certified a class of rightsholders whose books came from the pirated libraries on July 17, 2025, and set a piracy damages trial for December 2025. Statutory damages in a case with this many works run high: the theoretical exposure topped $70 billion, using the standard per-work statutory range.

Training on lawfully acquired books was fair use. Downloading and keeping pirated copies was not.— Summary judgment order, N.D. Cal., June 23, 2025

The number

Anthropic settled before that trial. On September 5, 2025, the parties disclosed a $1.5 billion deal covering about 482,460 works, plus destruction of the pirated files. It is the largest copyright settlement on record — larger than any prior music, publishing or software case.

The court granted preliminary approval later that month and opened a claims process. By the March 2026 deadline, class counsel reported claims covering more than 90 percent of eligible works — an unusually high participation rate for a class this size.

Judge Araceli Martínez-Olguín granted final approval on July 20, 2026.

$1.5 billion — about $3,000 per book — and expressly no binding precedent.— Order granting final approval, N.D. Cal., July 20, 2026

The objection the court itself raised

Class counsel asked for $187.5 million in attorneys’ fees. Martínez-Olguín cut that to about $101.6 million in the same order — a roughly 46 percent reduction, and the clearest signal in the ruling that the court was not simply rubber-stamping the deal.

That fee cut is the closest thing to a dissent inside an otherwise uncontested settlement. It says the court read the numbers closely enough to push back on the one figure that wasn’t fixed by negotiation with the defendant.

What the settlement doesn’t do

The order is explicit that it sets no binding precedent. It releases Anthropic’s past conduct for the class that opted in. It does not license anything going forward, and it does not cover claims about what Claude outputs — only what went into training it.

For data buyers, that leaves two separate numbers to track. Alsup’s fair-use ruling on purchased, scanned books is a district-court opinion, not settled law, but it is the closest thing the industry has to a green light for buy-then-train. The $3,000-per-book settlement figure is the reference price for the alternative: training on a shadow library and hoping nobody notices.

Why it matters for data buyers

Every other pending case in this tracker is arguing over the same two questions Bartz already answered for one defendant: does training on the work at all infringe, and does it matter how the copy was obtained? Bartz says provenance is the variable that decides outcomes, not the training itself.

A buyer diligencing a vendor’s dataset now has a concrete number to anchor a risk conversation: $3,000 per book is what unlicensed acquisition cost one company, after the fact, with no admission of liability and no forward license. Licensed acquisition remains the cheaper — and the only precedent-proof — path.

What to watch

  • Whether other pending author and publisher suits against AI developers — Kadrey v. Meta and Authors Guild v. OpenAI among them — cite Alsup’s fair-use/piracy split, even though it isn’t binding outside this case.
  • Whether the $3,000-per-work figure becomes a reference point in other shadow-library settlement talks, the way it’s already surfacing in commentary on the Suno and Udio music cases.
  • How the claims administrator distributes the $1.5 billion across roughly 482,460 works now that the claims window has closed.
  • Whether any class member opted out and pursues an individual damages claim now that final approval has been granted.

Settlement economics

~$1.5 billion — the largest copyright settlement in U.S. history, or about $3,000 per work.

Anthropic agreed to pay about $1.5 billion to rightsholders in roughly 482,460 books it downloaded from the LibGen and PiLiMi pirate libraries — about $3,000 per covered work — and to destroy the pirated files. Judge Araceli Martínez-Olguín granted final approval on July 20, 2026, and cut class counsel’s fee to about $101.6 million from the $187.5 million requested. Claims covered roughly 93 percent of eligible works.

The number came from leverage, not a market rate. Judge Alsup had ruled that training on lawfully bought-and-scanned books was fair use, but that downloading and keeping pirated copies was not, and he certified a class facing a December 2025 statutory-damages trial with theoretical exposure above $70 billion. About $3,000 per work is what that piracy exposure settled for — a release for past conduct, with no forward license, no output coverage, and, as the court stressed, no binding precedent.

Licensing lens. This prices unlicensed shadow-library sourcing, not licensed data: about $3,000 per book is now the reference cost of having pirated a work, while lawful acquisition sat on the fair-use side of the line. For buyers and sellers, provenance — how the data was obtained — is what carries the price, and a settlement buys peace for the past without licensing the future.

Sources

AnthropicCopyright settlementFair useBooksShadow libraries

See something wrong? Send a correction.

Jaeden Schafer

Jaeden Schafer

Jaeden Schafer is the founder of fiund and host of the AI Chat podcast. He covers the training-data market and the lawsuits shaping it.

Want data that clears this in diligence?

Whether you're building a model or sitting on an archive, the first conversation is short and specific.

Send a brief