Lawsuit tracker
U.S. District Court for the Southern District of New York (Judge Sidney H. Stein) · ActiveAuthors Guild v. OpenAI
Seventeen novelists and the Authors Guild turned a claim that ChatGPT was trained on their books into the largest consolidated author action against a model developer — now entangled with a fight over a training dataset OpenAI deleted years before the suit.
Key facts
- Filed September 19, 2023 in S.D.N.Y. by the Authors Guild and seventeen authors, including John Grisham, George R.R. Martin and Jodi Picoult; Microsoft was added as a defendant that December.
- Centralized in April 2025 into MDL No. 3143 before Judge Sidney Stein, alongside the Tremblay, Silverman and New York Times cases.
- In October 2025, Judge Stein denied OpenAI’s motion to dismiss the authors’ output-based direct-infringement claim.
- Discovery surfaced a fight over OpenAI’s 2022 deletion of two LibGen-derived training datasets, with a magistrate judge ordering production of internal Slack messages about the deletion.
- Fact discovery closed February 27, 2026; summary judgment reply briefs are now due November 6, 2026.
What the authors allege
Books are an easy claim to state and a hard one to prove: the Authors Guild and its seventeen named authors say their novels and nonfiction were copied, without permission or payment, into the datasets used to train GPT models. The theory covers both training — the copying itself — and outputs, arguing ChatGPT can reproduce or closely paraphrase passages from their books. Microsoft was added as a defendant in December 2023 for distributing and profiting from the models through its own products.
Consolidation into the OpenAI MDL
This case did not stay alone for long. In April 2025, the Judicial Panel on Multidistrict Litigation folded a dozen OpenAI copyright suits — including the earlier Tremblay and Silverman author cases from California and the New York Times’ suit — into MDL No. 3143 before Judge Sidney Stein in Manhattan. Consolidation does not merge the claims into one lawsuit; it centralizes pretrial proceedings like discovery so overlapping cases do not duplicate the same fights in different courtrooms. The author plaintiffs then filed a single consolidated class-action complaint in June 2025.
The output-infringement ruling
In October 2025, Stein denied OpenAI’s motion to dismiss the authors’ claim that ChatGPT’s outputs themselves — not just the training process — can infringe their copyrights. A motion to dismiss tests only whether the allegations, if true, could state a legal claim; the ruling means the authors plausibly alleged outputs can be substantially similar to their books, not that any specific output has been proven infringing. It keeps output-side liability alive as a track separate from the training-side theory that dominates most other AI copyright suits.
The fight over deleted training data
Discovery turned up a specific flashpoint: OpenAI deleted two training datasets, referred to as books1 and books2 and believed to be drawn from LibGen and other shadow libraries, back in 2022. Plaintiffs argue the deletion looks designed to obscure what the models were trained on; OpenAI has resisted producing related records. A magistrate judge ordered OpenAI in late 2025 to hand over internal Slack messages discussing the deletions — a ruling that, if the messages show intent to obscure infringing sourcing, could expose OpenAI to enhanced, willfulness-based damages later in the case.
Where it stands
Fact discovery closed on February 27, 2026. The court has since pushed the summary judgment schedule, with reply briefs now due November 6, 2026 rather than the original October date — a sign of how much material the parties are still working through. No ruling on the merits of infringement, fair use, or class certification has issued.
Why it matters for training-data licensing
This is the largest author class action in the space, and the deleted-dataset fight is the part that should worry any data buyer: a decision to delete source records, made years before litigation, is now itself potential evidence of consciousness of wrongdoing. For companies assembling training corpora today, the lesson is blunt — keep the receipts. Provenance records that look irrelevant at collection time can become the difference between an ordinary fair-use defense and a spoliation problem in front of a jury.
What to watch
- Summary judgment briefing, with replies due November 6, 2026.
- Whether the Slack messages about the deleted datasets support findings of willfulness or spoliation.
- Any class certification motion following the close of fact discovery.
- Whether the output-infringement theory becomes a template for other author suits inside the MDL.
Sources
- CourtListener docket: In re OpenAI, Inc. Copyright Infringement Litigation (MDL 3143)
- Saul Ewing: class of authors can pursue infringement claims against OpenAI
- The Hollywood Reporter: OpenAI loses discovery battle over deleted pirated-book datasets
- Chat GPT Is Eating the World: summary judgment briefing pushed to November 2026
- Authors Guild: understanding the AI class action lawsuits
See something wrong? Send a correction.
More case files
Want data that clears this in diligence?
Whether you're building a model or sitting on an archive, the first conversation is short and specific.
Send a brief