Lawsuit tracker
Publishers v. Google (Gemini)
Hachette Book Group, Inc., et al. v. Google LLC, No. 1:26-cv-05870 (S.D.N.Y.)
| Plaintiffs | Hachette Book Group, Cengage Learning, Elsevier, author Scott Turow and the group S.C.R.I.B.E., on behalf of a proposed class of book and journal copyright owners |
|---|---|
| Defendants | |
| Court | U.S. District Court, Southern District of New York |
| Filed | 2026-07-10 |
| Status | Active |
| Content type | books & journals |
| Last updated | 2026-07-21 |
The claims
Direct and contributory copyright infringement plus DMCA claims for removal or alteration of copyright-management information, which plaintiffs say was done to conceal that Gemini was trained on their works. The proposed class seeks statutory damages, an injunction and destruction of unauthorized training copies.
What has happened
Trade, education and academic publishers — joined by novelist Scott Turow and the group S.C.R.I.B.E. — filed a proposed class action alleging Google trained its Gemini models on millions of copyrighted books and journal articles without permission. The twist is the relationship history: plaintiffs had supplied works to Google for decades under scope-limited programs — Google Books scanning for search snippets, Google Play distribution — and allege Google reached into those corpora, plus pirate sites, for AI training it was never authorized to do. The complaint quotes an internal Google document allegedly warning that using copyrighted books for AI training could be “highly problematic” and bring “$10Bs-$100Bs in potential fines.” The filing was publicized by the Association of American Publishers on July 14, 2026; Google did not immediately comment.
Key developments
- 2026-07-10 — The class complaint is filed in the S.D.N.Y., pleading direct and contributory infringement and CMI-removal claims and seeking statutory damages, an injunction and destruction of unauthorized training copies.
- 2026-07-14 — The Association of American Publishers announces the case. Coverage notes it lands after 2025 California rulings that treated training as fair use but punished pirated sourcing — including the $1.5 billion Anthropic books settlement.
Why it matters for training data
A license scoped to indexing and snippets is not a license to train — legacy digitization and distribution deals are being re-read as liability. Google Books survived fair-use scrutiny for search; plaintiffs argue training is a different use with a different market, and they chose the S.D.N.Y. rather than the California courts that blessed training in 2025. For buyers and suppliers alike: scope clauses in old content agreements now carry model-training stakes, and internal risk memos are discoverable.
Sources
- TechCrunch: Google faces another AI training lawsuit
- AAP complaint PDF (Hachette v. Google, Dkt. 1)
- International Publishers Association announcement
- Al Jazeera: authors, publishers sue Google
Deeper analysis
Want data that clears this in diligence?
Whether you're building a model or sitting on an archive, the first conversation is short and specific.
Send a brief