Lawsuit tracker
S.D.N.Y. · ActivePublishers v. Google (Gemini)
Hachette, Cengage, Elsevier and novelist Scott Turow just accused Google of repurposing decades-old, scope-limited book deals to train Gemini. The case is days old: no answer from Google, no schedule, no ruling of any kind yet.
Key facts
- Hachette Book Group, Cengage Learning, Elsevier, author Scott Turow and the group S.C.R.I.B.E. filed a proposed class action against Google in the S.D.N.Y. on July 10, 2026 (No. 1:26-cv-05870).
- The complaint alleges Google trained its Gemini models on millions of copyrighted books and journal articles obtained under older, scope-limited programs — Google Books scanning for search snippets, and Google Play distribution — plus material from pirate sites.
- Claims include direct and contributory copyright infringement and DMCA claims for removing or altering copyright-management information, which plaintiffs say Google did to conceal its training sources.
- The complaint quotes an internal Google document allegedly warning that using copyrighted books for AI training could be "highly problematic" and risk "$10Bs-$100Bs in potential fines."
- As of July 22, 2026, no response from Google appears on the docket and no schedule has been set; this is a newly filed case with no ruling of any kind yet.
The allegations
The complaint, publicized by the Association of American Publishers on July 14, 2026, brings together trade, education and academic publishers with novelist Scott Turow and the author group S.C.R.I.B.E. Their proposed class covers book and journal copyright owners whose work they say Google fed into Gemini without authorization. The pleaded claims are direct and contributory copyright infringement, plus DMCA claims for removing or altering copyright-management information — the same CMI theory that has featured in several other AI publisher suits, here pleaded alongside a full infringement claim rather than standing alone. The proposed class seeks statutory damages, an injunction, and destruction of unauthorized training copies, across what the complaint frames as millions of individual works — a scale that, if a class is ever certified, would test how courts manage class-wide damages calculations for AI training claims at a size not yet seen in this wave of litigation.
The legal question
The distinguishing feature of this case is the relationship history between the parties. Plaintiffs say they supplied works to Google for decades under programs with a specific, limited scope: Google Books scanning for search indexing and snippet display, and Google Play for retail distribution. The suit alleges Google then used those same corpora — along with pirated sources — to train Gemini, a use the original agreements never authorized. The legal question is whether a license or program scoped to search and snippets, or to retail distribution, extends to training a generative model at all. That's a materially different question from suits alleging works were scraped from the open web with no prior relationship whatsoever.
The complaint also leans on an internal document it says shows Google's own lawyers flagged the risk before the fact, quoting language warning that training on copyrighted books could be "highly problematic" and could expose the company to "$10Bs-$100Bs in potential fines." That's a notable allegation — an internal risk assessment allegedly anticipating the very claim now being litigated — though it is, at this stage, only an allegation drawn from the complaint, not a document a court has authenticated or ruled on.
Forum choice and the shadow of 2025’s rulings
The plaintiffs filed in the Southern District of New York rather than the Northern District of California, where 2025 rulings — including the roughly $1.5 billion Anthropic books settlement — dealt with training-related copyright claims on more favorable terms for AI developers in some respects. Google Books itself survived a long-running fair-use challenge over search and snippets more than a decade ago. Plaintiffs are betting that training a generative model is a different use, in a different market, from indexing books for search — and that a different court, without that precedent directly on the books, is the better venue to test it.
Where it stands
This is a thin-fact case by necessity: it was filed July 10, 2026, and publicized July 14. As of July 22, 2026, there is no response from Google on the docket, no scheduling order, and no ruling — not on a motion to dismiss, not on class certification, not on anything. Everything about how this case will unfold, including whether it gets consolidated with other pending AI publisher litigation, remains open.
Why it matters for training-data licensing
This case is a direct test of an assumption a lot of buyers make: that a decades-old content relationship with a rights holder covers AI training just because a relationship exists. The plaintiffs' theory is that it doesn't — that scope language written for search indexing or retail distribution doesn't stretch to cover a new, different use invented years later. Anyone relying on a legacy digitization, syndication or distribution agreement as cover for AI training should read the actual grant language, not just confirm that a deal is in place.
The alleged internal risk memo is also worth watching for a separate reason: it suggests that a company's own contemporaneous legal analysis of its training-data risk can become the plaintiffs' best exhibit once litigation starts. That's a discoverability lesson as much as a licensing one, and it's consistent with what's played out in the parallel OpenAI and Cohere cases.
What to watch
- Google's initial response — an answer or a motion to dismiss — and the theories it raises.
- Whether the case is consolidated with, or coordinated alongside, other pending AI publisher litigation.
- How the scope-of-license theory fares against Google's likely defenses drawing on its earlier Google Books fair-use win.
- Any motion practice testing the internal-document allegations before discovery begins.
- Class-certification proceedings, once the case reaches that stage.
Sources
See something wrong? Send a correction.
More case files
Want data that clears this in diligence?
Whether you're building a model or sitting on an archive, the first conversation is short and specific.
Send a brief