Lawsuit tracker

U.S. District Court, Southern District of New York (Judge Sidney H. Stein) · Active

CIR v. OpenAI

A nonprofit newsroom’s own reading of OpenAI’s published training-data list gave it the evidence to sue — a reminder that a modest catalog can sustain years of federal litigation against a far larger opponent.

Key facts

  1. Filed June 27, 2024 in S.D.N.Y. by the Center for Investigative Reporting, publisher of Mother Jones and Reveal, against OpenAI and Microsoft.
  2. The complaint cites OpenAI’s own published list of top domains in its WebText training set, which ranked motherjones.com at No. 267 with 16,793 distinct URLs ingested.
  3. Claims copyright infringement and DMCA copyright-management-information violations over training use and ChatGPT outputs.
  4. Consolidated with the New York Times and Daily News cases before Judge Sidney Stein, and now proceeds inside MDL No. 3143.
  5. On July 9, 2026, CIR joined the Times, Daily News, Intercept and Ziff Davis in a sanctions motion alleging OpenAI withheld and destroyed evidence; the motion is pending.
PartiesThe Center for Investigative Reporting, publisher of Mother Jones and Reveal v. OpenAI entities and Microsoft
CourtU.S. District Court, Southern District of New York (Judge Sidney H. Stein)
Docket3143
Filed2024-06-27
Content typejournalism
StatusAs of July 2026, active. CIR’s claims proceed in consolidated pretrial proceedings with the Times and Daily News cases inside the OpenAI MDL.

How a domain ranking became a complaint

The Center for Investigative Reporting’s suit stands out for its evidentiary shortcut: rather than relying on inference about what might have been scraped, CIR pointed to a domain list OpenAI itself published describing sources behind its WebText training corpus. Motherjones.com ranked 267th on that list, with 16,793 distinct URLs pulled in — a specific, citable figure that let CIR allege ingestion directly rather than argue from circumstantial similarity between outputs and its stories.

The claims

CIR’s complaint pairs the now-familiar copyright infringement theory — that its journalism was copied into training data without permission or payment — with a DMCA claim over copyright-management information, arguing that stripping bylines, credits or copyright notices during scraping and training violates a statutory protection separate from the infringement claim itself. The nonprofit argues other organizations pay to license the same reporting that trained ChatGPT for free.

Folded into the OpenAI MDL

CIR’s case was consolidated with the New York Times and Daily News actions before Judge Sidney Stein and now shares the same pretrial discovery track as those larger newspaper plaintiffs inside MDL No. 3143. That means CIR benefits from discovery fights fought and won by better-resourced co-plaintiffs — including the fight over OpenAI’s chat-log production — without having to litigate each dispute independently.

The sanctions fight

On July 9, 2026, CIR joined the Times, Daily News, Intercept and Ziff Davis in moving for sanctions, alleging OpenAI withheld and destroyed evidence relevant to the litigation. Sanctions motions ask a court to penalize a party for its conduct during the case itself — separate from the underlying merits of infringement — and can range from adverse-inference instructions to monetary penalties, depending on what the court finds. The motion is pending.

Why it matters for training-data licensing

CIR is proof that scale does not determine litigation risk: a single-newsroom nonprofit, on contingency-fee counsel, has sustained a federal case against two of the best-resourced defendants in tech for two years and counting, on the strength of a training-data disclosure the defendant itself published. For any company assembling a training corpus, publishing detailed dataset manifests is a double-edged transparency move — it builds trust with regulators and partners, but it also hands future plaintiffs their proof of ingestion. Registered works with even modest reach carry the same legal exposure as marquee mastheads once ingestion can be shown.

What to watch

  • The pending sanctions motion over OpenAI’s alleged withholding or destruction of evidence.
  • Coordination of discovery rulings across the consolidated newspaper and nonprofit plaintiffs in the MDL.
  • Whether CIR’s DMCA copyright-management-information claim survives as a track separate from the core infringement count.
  • Any bellwether-style ruling in the MDL that could set terms for smaller publisher claims.

Sources

copyrightjournalismDMCAdiscoverytraining data

See something wrong? Send a correction.

Jaeden Schafer

Jaeden Schafer

Jaeden Schafer is the founder of fiund and host of the AI Chat podcast. He covers the training-data market and the lawsuits shaping it.

Want data that clears this in diligence?

Whether you're building a model or sitting on an archive, the first conversation is short and specific.

Send a brief