Lawsuit tracker
MDL No. 3143 · ActiveNYT v. OpenAI
The Times' case against OpenAI and Microsoft is deep in discovery inside a federal multidistrict litigation — with a 20-million-log production order and a pending sanctions motion — but no court has ruled on fair use, and there is no trial date.
Key facts
- The New York Times sued OpenAI and Microsoft in the S.D.N.Y. on December 27, 2023; the case is now part of In re: OpenAI, Inc. Copyright Infringement Litigation (MDL No. 3143) before Judge Sidney Stein.
- Judge Stein largely denied the motions to dismiss in a March 26, 2025 ruling (opinion issued April 4, 2025), keeping the core copyright claims — including contributory infringement — in the case.
- On January 5, 2026, Stein affirmed a magistrate judge's order requiring OpenAI to produce a sample of roughly 20 million de-identified ChatGPT conversation logs.
- On July 9, 2026, the Times, joined by the Daily News, the Center for Investigative Reporting, The Intercept and Ziff Davis, moved for sanctions, alleging OpenAI concealed an internal database of about 78 million de-identified chats and deleted logs subject to preservation orders; the motion is pending.
- No court has ruled on the merits of OpenAI's fair-use defense, and no trial date has been set.
The allegations
The Times was the first major American newspaper to sue over generative AI, filing on December 27, 2023 against both OpenAI and its primary financial backer, Microsoft. The complaint alleges direct and contributory copyright infringement, DMCA copyright-management-information violations, unfair competition and trademark dilution, and it attached roughly one hundred exhibits showing ChatGPT reproducing Times journalism nearly verbatim. The Times has put its damages claim in the billions of dollars and asked the court to order destruction of the models and training sets built on its work.
OpenAI's defense, consistent across the MDL, is that training on publicly available text is transformative fair use, and that the Times engineered its verbatim-output exhibits through contrived, adversarial prompting that doesn't reflect ordinary use.
What the court has — and hasn’t — decided
Judge Stein let the core copyright claims through in his March 2025 ruling on the motions to dismiss, including the contributory-infringement theory, while trimming some peripheral claims. That ruling matters, but it is not a decision on fair use. A motion to dismiss asks only whether the plaintiff has pled a legally viable claim; surviving one means the case gets to proceed, not that the Times has won or that OpenAI's fair-use defense has failed. No fair-use ruling, on summary judgment or otherwise, has issued in this case.
Shortly after, the case was consolidated into MDL 3143, gathering the federal OpenAI copyright suits — other newspapers, magazine publishers, authors' groups and the New York Daily News coalition among them — for coordinated pretrial proceedings before Judge Stein, with discovery supervised by Magistrate Judge Ona Wang. That consolidation is part of why discovery rulings in the Times' case, like the 20-million-log production order, tend to apply across the MDL rather than to a single plaintiff, and why a single fair-use ruling here would likely shape the parallel publisher suits against OpenAI as well.
Discovery as the real battleground
Since consolidation, the fight has largely been about what OpenAI must turn over. In January 2026, Stein affirmed Magistrate Wang's order requiring production of roughly 20 million de-identified ChatGPT conversation logs, rejecting OpenAI's user-privacy objections and holding the logs bear on both output behavior and the fair-use defense. In spring 2026, after finding an OpenAI corporate witness unprepared on noticed deposition topics, Wang ordered a properly prepared witness produced for re-deposition and gave plaintiffs added deposition time.
The dispute escalated further on July 9, 2026, when the Times and four co-plaintiffs moved for sanctions. The motion alleges OpenAI had already searched its own systems for copyrighted journalism and had assembled an internal database of roughly 78 million de-identified chats before the litigation began — while representing to the court that no such search was feasible — and that it deleted logs covered by preservation orders. Plaintiffs are asking the court to bar OpenAI from relying on its own reduced log sample and to treat as established fact that the full logs would show reproduction of their journalism. OpenAI has denied wrongdoing and says the motion is an attempt to paper over a weakening case. The motion remains undecided.
Where it stands
As of July 22, 2026, the case is active within the MDL, in discovery, with no trial date. The sanctions motion is pending, and the underlying merits question — whether training ChatGPT on Times journalism was fair use — has not been decided by any court.
Why it matters for training-data licensing
This case is setting the discovery playbook for AI copyright litigation generally: usage logs, training-corpus audits and executives' internal notes have all proven discoverable, and how a company retains or deletes that material has become litigation infrastructure in its own right. The sanctions fight shows courts are willing to police that conduct directly.
For data buyers and suppliers, the eventual fair-use ruling here — whenever it comes — will effectively set a price signal for news content across the industry, since it's the most closely watched of the parallel publisher suits against major model builders. In the meantime, the discovery record is the more immediate lesson: retention policies, deletion practices and internal risk assessments about training data are not back-office matters. They are evidence.
What to watch
- A ruling on the pending July 2026 sanctions motion.
- What the 20-million-log production ultimately shows once reviewed.
- Whether the case reaches summary judgment on fair use, or settles first, as some other AI publisher disputes have.
- Coordination between this MDL and parallel publisher suits against Cohere and Google over similar training-data theories.
Sources
See something wrong? Send a correction.
More case files
Want data that clears this in diligence?
Whether you're building a model or sitting on an archive, the first conversation is short and specific.
Send a brief