Lawsuit tracker
U.S. District Court, Southern District of New York (Judge Sidney H. Stein) · ActiveDaily News v. OpenAI
Eight Alden-owned dailies sued OpenAI and Microsoft together, arguing their reporting was copied into training data and now resurfaces — sometimes fabricated — in ChatGPT and Copilot answers.
Key facts
- Filed April 30, 2024 in S.D.N.Y. by eight newspapers under Alden Global Capital’s MediaNews Group and Tribune Publishing, including the New York Daily News, Chicago Tribune, Orlando Sentinel, Denver Post and San Jose Mercury News.
- Alleges copyright infringement over training and outputs, DMCA violations, and that models fabricate content falsely attributed to the papers’ mastheads.
- Consolidated with the New York Times case before Judge Sidney Stein; core copyright claims survived dismissal motions on March 26, 2025.
- On January 5, 2026, Judge Stein affirmed an order requiring OpenAI to produce roughly 20 million de-identified ChatGPT logs to the newspaper plaintiffs.
- On July 9, 2026, the Daily News plaintiffs joined other outlets in a pending sanctions motion over OpenAI’s alleged discovery conduct.
Eight papers, one financial owner
Unlike most AI-training suits filed by a single publisher, this case bundles eight metro dailies under one owner: Alden Global Capital, through its MediaNews Group and Tribune Publishing units. Alden has a reputation for treating local newspapers primarily as financial assets managed for cash flow — and this suit extends that logic to the papers’ archives, arguing decades of regional reporting are licensable property whether or not the newsrooms that produced it are shrinking.
What the complaint alleges
The claims track the now-standard newspaper theory against OpenAI: that copyrighted articles were copied into training datasets without permission, that ChatGPT and Microsoft’s Copilot can reproduce or closely paraphrase that reporting in outputs, and that copyright-management information was stripped in the process, supporting a separate DMCA claim. The papers add a reputational layer beyond the New York Times case’s core theory — that the models sometimes fabricate stories or quotes and attribute them to the papers’ mastheads, potentially harming publications for reporting they never did.
Surviving dismissal, then discovery
The case was consolidated with the Times’ suit before Judge Sidney Stein, and on March 26, 2025, Stein turned back the defendants’ dismissal motions across the consolidated newspaper cases, letting the core copyright claims proceed. A motion-to-dismiss survival tests only whether the complaint states a plausible claim — it is not a ruling that infringement occurred. Discovery has since produced one of the largest data-production fights in the litigation: on January 5, 2026, Stein affirmed an order requiring OpenAI to hand over roughly 20 million de-identified ChatGPT conversation logs to the newspaper plaintiffs, material the papers say is needed to show how often and how closely the models reproduce their reporting in real user interactions.
The sanctions motion
On July 9, 2026, the Daily News plaintiffs joined the Times and other newspaper and magazine plaintiffs in the MDL in moving for sanctions, alleging OpenAI withheld or destroyed evidence relevant to the case. That motion is separate from, and additional to, the underlying infringement claims, and it remains pending.
Why it matters for training-data licensing
This case is a useful corrective to the idea that AI-training litigation risk tracks a publisher’s fame: Alden’s mid-market titles carry the same legal exposure as the country’s most prominent masthead once a registered archive and a scraped training run are in the picture. For buyers assessing a dataset or a model’s supply chain, that means checking exposure across regional and financial-owner-held publications, not just marquee names — one MDL ruling on the shared discovery record could reprice licensing risk for an entire tier of news content at once.
What to watch
- The pending sanctions motion over OpenAI’s discovery conduct.
- How the 20-million-log production is used to argue output-based infringement.
- Whether the fabrication and misattribution allegations develop into a standalone theory distinct from copying claims.
- Coordination with the broader MDL’s summary judgment schedule.
Sources
See something wrong? Send a correction.
More case files
Want data that clears this in diligence?
Whether you're building a model or sitting on an archive, the first conversation is short and specific.
Send a brief