Lawsuit tracker

U.S. District Court for the Northern District of California, San Jose (Judge Eumi K. Lee) · Active

Zhang v. Google

Four visual artists built their case on a dataset Google itself disclosed in a research paper. LAION-400M did the plaintiffs' tracing work for them — but a year and a half in, the fight over whether they've identified enough infringed works to proceed is still not resolved.

Key facts

  1. Zhang v. Google LLC, No. 5:24-cv-02531 (N.D. Cal.), filed April 26, 2024 by photographer Jingna Zhang and cartoonists Sarah Andersen, Hope Larson and Jessica Fink on behalf of a proposed class of visual artists.
  2. The complaint alleges Google's Imagen text-to-image model was trained on the LAION-400M dataset, which Google itself had disclosed, and that the dataset included the plaintiffs' registered works.
  3. The case was consolidated in late 2024 into In re Google Generative AI Copyright Litigation, merging the artists’ claims with author Jill Leovy’s separate suit; a consolidated amended complaint was filed December 20, 2024.
  4. Google moved to dismiss parts of the consolidated complaint in January 2025, arguing plaintiffs had not identified the specific works infringed; Judge Eumi K. Lee heard that motion on April 23, 2025 and took it under submission after signaling she was inclined to dismiss some claims.
  5. Proceedings continued into 2026, including a February 20, 2026 hearing, with the consolidated case now in discovery.
PartiesJingna Zhang, Sarah Andersen, Hope Larson and Jessica Fink, on behalf of a proposed class of visual artists v. Google LLC and Alphabet Inc.
CourtU.S. District Court for the Northern District of California, San Jose (Judge Eumi K. Lee)
Docket5:24-cv-02531
Filed2024-04-26
Content typeimages
StatusAs of July 2026, active as part of the consolidated In re Google Generative AI Copyright Litigation; the case is in discovery, with Google’s April 2025 motion to dismiss taken under submission.

A dataset Google named itself

Photographer Jingna Zhang and cartoonists Sarah Andersen, Hope Larson and Jessica Fink sued Google in April 2024, alleging its Imagen text-to-image model was trained on LAION-400M — a large open dataset of image-text pairs. The detail that made the case possible: Google had disclosed that dataset by name in its own published research on Imagen.

That disclosure did the plaintiffs’ tracing work for them. Rather than needing to reverse-engineer what a closed model was trained on, they could point to LAION-400M’s known contents and argue their registered images were in it.

Consolidation with the Leovy case

In late 2024, the court folded the artists’ suit into In re Google Generative AI Copyright Litigation alongside a separate suit brought by author Jill Leovy, combining visual-art and text claims against Google under one consolidated proceeding. A consolidated amended complaint was filed December 20, 2024.

Google moved to dismiss parts of that consolidated complaint in January 2025, arguing the plaintiffs had not specifically identified which of their works were infringed — a threshold pleading argument common in AI training cases, aimed at the complaint’s specificity rather than the merits of fair use.

A signal, not a ruling

Judge Eumi K. Lee heard Google's motion on April 23, 2025 and indicated she was inclined to dismiss some of the copyright claims, then took the matter under submission rather than ruling from the bench. That is a signal about the judge's thinking, not a decision — no order dismissing any claims had been reported as of that hearing, and it's a mistake to treat a judge's questioning at oral argument as a ruling.

Proceedings continued through 2026, including a further hearing on February 20, 2026, with the consolidated case moving into discovery even as the specificity questions from the motion to dismiss remain part of the ongoing proceedings.

Why the LAION disclosure is the real story

The lasting significance of this case may have less to do with its outcome than with how it started. Google’s own published paper, naming its training dataset, became the evidentiary foundation for a copyright suit against Google. That is a lesson every model developer should sit with: a dataset citation in a research paper is a discoverable, quotable admission, not a footnote that disappears once the paper is published.

For data suppliers, the read-across is just as direct. Open, web-scraped datasets like LAION carry the same infringement exposure as any proprietary scrape — arguably more, since their contents are publicly documented and searchable by the very people whose works might be in them. "It’s an open academic dataset" is not a chain-of-title defense; it may be the opposite, a public map of exactly what to check for.

What to watch

  • Judge Lee’s ruling on Google’s motion to dismiss, still under submission as of the last reported hearing.
  • How specifically plaintiffs must identify individual infringed works to survive a motion to dismiss in an AI training case — a question with implications well beyond this suit.
  • Progress in discovery within the consolidated In re Google Generative AI Copyright Litigation proceeding.
  • Whether other artists whose work appears in LAION-400M file similar suits against other developers who trained on the same public dataset.

Sources

GoogleImagenLAIONVisual artistsImage generation

See something wrong? Send a correction.

Jaeden Schafer

Jaeden Schafer

Jaeden Schafer is the founder of fiund and host of the AI Chat podcast. He covers the training-data market and the lawsuits shaping it.

Want data that clears this in diligence?

Whether you're building a model or sitting on an archive, the first conversation is short and specific.

Send a brief