Lawsuit tracker

N.D. Cal. · Active

Nazemian v. NVIDIA

A federal judge let novelists' piracy claims against NVIDIA's NeMo Megatron models clear the pleadings stage, on a theory that reaches the tooling used to assemble a training set, not just the party running the training job.

Key facts

  1. Novelists Abdi Nazemian, Brian Keene and Stewart O'Nan sued NVIDIA in the N.D. Cal. on March 8, 2024; a related suit by Andre Dubus III and Susan Orlean was later consolidated into the case.
  2. The suit alleges NVIDIA trained its NeMo Megatron large language models on The Pile, including the Books3 collection of nearly 200,000 books copied from the shadow library Bibliotik.
  3. In October 2025, plaintiffs won leave to amend their complaint to add allegations involving additional shadow libraries, including claims that NVIDIA contacted the shadow library Anna's Archive seeking access to pirated books.
  4. On May 5, 2026, Judge Jon Tigar denied NVIDIA's motion to dismiss the amended complaint as to direct and contributory copyright infringement, rejecting NVIDIA's comparison of itself to a passive internet service provider.
  5. The court dismissed a vicarious-infringement claim, with leave to amend, for failing to allege NVIDIA controlled or profited directly from the underlying infringement.
  6. The ruling is a pleadings-stage decision only; there is no finding of infringement, no fair-use ruling and no trial date.
PartiesAuthors Abdi Nazemian, Brian Keene and Stewart O’Nan, later joined by Andre Dubus III and Susan Orlean, on behalf of a proposed class v. NVIDIA Corporation
CourtU.S. District Court for the Northern District of California (Judge Jon S. Tigar)
Docket24-cv-01454
Filed2024-03-08
Content typeBooks — nearly 200,000 titles in the Books3 collection within The Pile, sourced from the shadow library Bibliotik
StatusAs of July 2026, active. On May 5, 2026, Judge Jon Tigar denied most of NVIDIA’s motion to dismiss, letting direct and contributory infringement claims proceed; discovery into NVIDIA’s shadow-library datasets is under way.

The allegations

The case began in March 2024 when three novelists accused NVIDIA of training its NeMo Megatron family of large language models on The Pile, an open-source training corpus assembled by EleutherAI. Buried inside The Pile is Books3, a collection of nearly 200,000 books copied wholesale from the shadow library Bibliotik. A related suit brought by Andre Dubus III and Susan Orlean was later folded into the case, adding their claims to the same theory.

The complaint pleads three variations of copyright liability: direct infringement, for copying the books into NVIDIA's training pipeline; contributory infringement, for allegedly providing tools that facilitated the underlying piracy; and vicarious infringement, for allegedly profiting from and having the ability to control that piracy.

In October 2025, before the motion-to-dismiss ruling, the plaintiffs sought and won leave to amend the complaint to widen its scope, adding allegations that NVIDIA drew on shadow libraries beyond Books3 and had contacted Anna's Archive, another well-known piracy repository, seeking access to millions of pirated books. NVIDIA moved to dismiss that amended complaint on January 29, 2026, arguing the plaintiffs still hadn't plausibly alleged which specific works were copied, when the copying occurred, or which of its models contained them. Judge Tigar's May 2026 ruling addressed that amended pleading.

The legal question

Direct infringement here doesn't turn on what NVIDIA's models produce when prompted — it turns on whether copying the books into a training dataset was itself an unauthorized reproduction, regardless of downstream output. Contributory infringement reaches further: the authors allege that scripts within NVIDIA's own NeMo framework existed to speed up downloading of pirated books, and that supplying that tooling can make NVIDIA liable for infringement it didn't personally commit.

The ruling and its limits

On May 5, 2026, Judge Tigar denied NVIDIA's motion to dismiss in large part. He held the authors plausibly alleged direct infringement from the copying of Books3 into NVIDIA's training corpus, and that the NeMo scripts could support contributory liability — expressly rejecting NVIDIA's argument that it should be treated like a passive host under the internet-service-provider framework that shields platforms from liability for what users upload. That rejection is itself notable doctrine: the ISP framework has shielded platforms hosting user uploads for decades, and Tigar's refusal to extend it to a company that built and distributed the training tools signals real skepticism toward infrastructure-level defenses generally, not just NVIDIA's version of it. The vicarious infringement claim did not survive: the court dismissed it, with leave to amend, for failing to allege NVIDIA had the right and ability to control the underlying infringing conduct or derived a direct financial benefit from it.

It's worth being precise about what this ruling is and isn't. It is a pleadings-stage decision holding the authors' theories are legally viable enough to proceed to discovery. It is not a finding that NVIDIA infringed anything, and it says nothing about whether training on Books3 would ultimately qualify as fair use — that defense hasn't been adjudicated in this case.

Where it stands

Post-motion-to-dismiss, the case is in discovery. In spring 2026 the court pressed NVIDIA to produce information about the shadow-library datasets used in training. No trial date has been set, and the authors have leave to file an amended vicarious-infringement claim.

Why it matters for training-data licensing

NVIDIA is best known as an infrastructure and tooling company, not a chatbot maker, and that's what makes this ruling notable: it extends training-data liability into the toolchain. A contributory-infringement theory built on code that allegedly sped up pirated downloads reaches vendors of data-preparation tooling, not only whoever runs the final training job.

It's also another major suit anchored on Books3, joining a growing list of cases that trace back to the same pirated corpus across multiple, unrelated defendants. For buyers, that pattern is the practical takeaway: any dataset with roots in Books3 or the Bibliotik shadow library carries litigation exposure wherever it resurfaces, and tooling vendors should now assume plaintiffs will read their code, not just evaluate their models' outputs. The addition of Anna's Archive to the complaint shows how these suits keep expanding as plaintiffs' lawyers trace the same infrastructure company's touchpoints with multiple shadow libraries, not just one.

What to watch

  • Whether the authors file — and the court accepts — an amended vicarious-infringement claim.
  • Discovery into NVIDIA's sourcing of shadow-library datasets, including the Anna's Archive allegations.
  • Any eventual ruling on NVIDIA's fair-use defense, which hasn't been reached yet.
  • Class-certification proceedings, and whether other infrastructure or tooling vendors face similar contributory-liability theories.

Sources

booksshadow librariescopyrightNVIDIAclass action

See something wrong? Send a correction.

Jaeden Schafer

Jaeden Schafer

Jaeden Schafer is the founder of fiund and host of the AI Chat podcast. He covers the training-data market and the lawsuits shaping it.

Want data that clears this in diligence?

Whether you're building a model or sitting on an archive, the first conversation is short and specific.

Send a brief