News
Training-data lawsuit news: week of July 20, 2026
Published 2026-07-21
Anthropic’s $1.5 billion authors’ settlement got final approval, and a new publisher class action against Google widened the plaintiff pool.
Anthropic’s $1.5B authors’ settlement got final approval
U.S. District Judge Araceli Martínez-Olguín granted final approval on July 20, 2026 — the largest known US copyright settlement, about $3,000 per book, and expressly no binding precedent. The 23-page order also cut class counsel’s fee award to roughly $101.5 million from the $187.5 million sought. It followed Judge Alsup’s June 2025 ruling that training on lawfully acquired books was fair use while retaining a pirated library was not.
Why it matters: The number prices unlicensed acquisition after the fact; licensed acquisition is the cheaper path.
Publishers filed a new class action against Google over Gemini training
The suit was filed July 10, 2026 in the Southern District of New York. Plaintiffs including Hachette, Cengage, Elsevier, author Scott Turow, and the group S.C.R.I.B.E. allege Google trained Gemini on their copyrighted works — including books supplied under scope-limited Google Books and Google Play Books agreements — and stripped copyright-management information to conceal it.
Why it matters: The plaintiff pool now includes education and scientific publishers, widening the exposed content classes beyond trade books and news.
Suno’s fair-use ruling slipped into 2027 as the labels expanded their claims
The record labels’ case against AI music generator Suno remained active in the District of Massachusetts before Chief Judge F. Dennis Saylor IV, with UMG and Sony pressing on after Warner’s November 2025 settlement. Audible Magic audio fingerprinting identified 61,026 copyrighted recordings in Suno’s training data — described by the labels as a small fraction of total matches — and on May 21, 2026 they moved to expand the complaint to cover them. A June 30, 2026 amended scheduling order on the docket ran fact discovery to September 30, 2026 and set dispositive motions for April 9, 2027, pushing any US fair-use ruling on music training into 2027.
Why it matters: Fingerprinting turned "what is in the training set" from an allegation into evidence — the same diligence question buyers should ask any audio vendor.
Still pending: the Third Circuit’s decision in Thomson Reuters v. Ross
Oral argument was heard June 11, 2026 — the first appellate test of fair use for AI training. The decision is pending, and it will bind more broadly than the non-generative district ruling.
Why it matters: The single biggest pending catalyst for how US law treats training on licensed rivals’ content.
NYT v. OpenAI grinds through discovery
As of mid-2026 the case remained in discovery, with no trial date set. OpenAI was ordered to produce roughly 20 million ChatGPT conversation logs — an order affirmed January 5, 2026 — and in April 2026 was ordered to re-produce an unprepared corporate designee. On July 9, 2026, the Times, the Daily News, and other news plaintiffs moved for sanctions, accusing OpenAI of concealing for over two years its ability to search its training datasets and output logs.
Why it matters: No merits ruling exists — any vendor or headline claiming a winner is wrong.
Training-data copyright lawsuits: the landscape ← All roundups
Want data that clears this in diligence?
Whether you're building a model or sitting on an archive, the first conversation is short and specific.
Send a brief