Lawsuit tracker
U.S. District Court for the Northern District of California, San Jose (Judge Eumi K. Lee) · ActiveZhang v. Google
Four visual artists built their case on a dataset Google itself disclosed in a research paper. LAION-400M did the plaintiffs' tracing work for them — but a year and a half in, the fight over whether they've identified enough infringed works to proceed is still not resolved.
Key facts
- Zhang v. Google LLC, No. 5:24-cv-02531 (N.D. Cal.), filed April 26, 2024 by photographer Jingna Zhang and cartoonists Sarah Andersen, Hope Larson and Jessica Fink on behalf of a proposed class of visual artists.
- The complaint alleges Google's Imagen text-to-image model was trained on the LAION-400M dataset, which Google itself had disclosed, and that the dataset included the plaintiffs' registered works.
- The case was consolidated in late 2024 into In re Google Generative AI Copyright Litigation, merging the artists’ claims with author Jill Leovy’s separate suit; a consolidated amended complaint was filed December 20, 2024.
- Google moved to dismiss parts of the consolidated complaint in January 2025, arguing plaintiffs had not identified the specific works infringed; Judge Eumi K. Lee heard that motion on April 23, 2025 and took it under submission after signaling she was inclined to dismiss some claims.
- Proceedings continued into 2026, including a February 20, 2026 hearing, with the consolidated case now in discovery.
A dataset Google named itself
Photographer Jingna Zhang and cartoonists Sarah Andersen, Hope Larson and Jessica Fink sued Google in April 2024, alleging its Imagen text-to-image model was trained on LAION-400M — a large open dataset of image-text pairs. The detail that made the case possible: Google had disclosed that dataset by name in its own published research on Imagen.
That disclosure did the plaintiffs’ tracing work for them. Rather than needing to reverse-engineer what a closed model was trained on, they could point to LAION-400M’s known contents and argue their registered images were in it.
Consolidation with the Leovy case
In late 2024, the court folded the artists’ suit into In re Google Generative AI Copyright Litigation alongside a separate suit brought by author Jill Leovy, combining visual-art and text claims against Google under one consolidated proceeding. A consolidated amended complaint was filed December 20, 2024.
Google moved to dismiss parts of that consolidated complaint in January 2025, arguing the plaintiffs had not specifically identified which of their works were infringed — a threshold pleading argument common in AI training cases, aimed at the complaint’s specificity rather than the merits of fair use.
A signal, not a ruling
Judge Eumi K. Lee heard Google's motion on April 23, 2025 and indicated she was inclined to dismiss some of the copyright claims, then took the matter under submission rather than ruling from the bench. That is a signal about the judge's thinking, not a decision — no order dismissing any claims had been reported as of that hearing, and it's a mistake to treat a judge's questioning at oral argument as a ruling.
Proceedings continued through 2026, including a further hearing on February 20, 2026, with the consolidated case moving into discovery even as the specificity questions from the motion to dismiss remain part of the ongoing proceedings.
Why the LAION disclosure is the real story
The lasting significance of this case may have less to do with its outcome than with how it started. Google’s own published paper, naming its training dataset, became the evidentiary foundation for a copyright suit against Google. That is a lesson every model developer should sit with: a dataset citation in a research paper is a discoverable, quotable admission, not a footnote that disappears once the paper is published.
For data suppliers, the read-across is just as direct. Open, web-scraped datasets like LAION carry the same infringement exposure as any proprietary scrape — arguably more, since their contents are publicly documented and searchable by the very people whose works might be in them. "It’s an open academic dataset" is not a chain-of-title defense; it may be the opposite, a public map of exactly what to check for.
What to watch
- Judge Lee’s ruling on Google’s motion to dismiss, still under submission as of the last reported hearing.
- How specifically plaintiffs must identify individual infringed works to survive a motion to dismiss in an AI training case — a question with implications well beyond this suit.
- Progress in discovery within the consolidated In re Google Generative AI Copyright Litigation proceeding.
- Whether other artists whose work appears in LAION-400M file similar suits against other developers who trained on the same public dataset.
Sources
See something wrong? Send a correction.
More case files
Want data that clears this in diligence?
Whether you're building a model or sitting on an archive, the first conversation is short and specific.
Send a brief