Lawsuit tracker

U.S. District Court for the Southern District of New York (Judge Sidney H. Stein) · Active

Authors Guild v. OpenAI

Seventeen novelists and the Authors Guild turned a claim that ChatGPT was trained on their books into the largest consolidated author action against a model developer — now entangled with a fight over a training dataset OpenAI deleted years before the suit.

Key facts

  1. Filed September 19, 2023 in S.D.N.Y. by the Authors Guild and seventeen authors, including John Grisham, George R.R. Martin and Jodi Picoult; Microsoft was added as a defendant that December.
  2. Centralized in April 2025 into MDL No. 3143 before Judge Sidney Stein, alongside the Tremblay, Silverman and New York Times cases.
  3. In October 2025, Judge Stein denied OpenAI’s motion to dismiss the authors’ output-based direct-infringement claim.
  4. Discovery surfaced a fight over OpenAI’s 2022 deletion of two LibGen-derived training datasets, with a magistrate judge ordering production of internal Slack messages about the deletion.
  5. Fact discovery closed February 27, 2026; summary judgment reply briefs are now due November 6, 2026.
PartiesThe Authors Guild and seventeen named authors, including John Grisham, George R.R. Martin, Jodi Picoult and Jonathan Franzen, on behalf of a proposed class v. OpenAI entities and Microsoft Corporation
CourtU.S. District Court for the Southern District of New York (Judge Sidney H. Stein)
Docket23-cv-08292
Filed2023-09-19
Content typeBooks — fiction and nonfiction, including works allegedly obtained from shadow-library datasets
StatusAs of July 2026, active in the SDNY multidistrict litigation. Fact discovery closed in February 2026 and summary judgment briefing is scheduled to run into November 2026; no fair-use ruling yet.

What the authors allege

Books are an easy claim to state and a hard one to prove: the Authors Guild and its seventeen named authors say their novels and nonfiction were copied, without permission or payment, into the datasets used to train GPT models. The theory covers both training — the copying itself — and outputs, arguing ChatGPT can reproduce or closely paraphrase passages from their books. Microsoft was added as a defendant in December 2023 for distributing and profiting from the models through its own products.

Consolidation into the OpenAI MDL

This case did not stay alone for long. In April 2025, the Judicial Panel on Multidistrict Litigation folded a dozen OpenAI copyright suits — including the earlier Tremblay and Silverman author cases from California and the New York Times’ suit — into MDL No. 3143 before Judge Sidney Stein in Manhattan. Consolidation does not merge the claims into one lawsuit; it centralizes pretrial proceedings like discovery so overlapping cases do not duplicate the same fights in different courtrooms. The author plaintiffs then filed a single consolidated class-action complaint in June 2025.

The output-infringement ruling

In October 2025, Stein denied OpenAI’s motion to dismiss the authors’ claim that ChatGPT’s outputs themselves — not just the training process — can infringe their copyrights. A motion to dismiss tests only whether the allegations, if true, could state a legal claim; the ruling means the authors plausibly alleged outputs can be substantially similar to their books, not that any specific output has been proven infringing. It keeps output-side liability alive as a track separate from the training-side theory that dominates most other AI copyright suits.

The fight over deleted training data

Discovery turned up a specific flashpoint: OpenAI deleted two training datasets, referred to as books1 and books2 and believed to be drawn from LibGen and other shadow libraries, back in 2022. Plaintiffs argue the deletion looks designed to obscure what the models were trained on; OpenAI has resisted producing related records. A magistrate judge ordered OpenAI in late 2025 to hand over internal Slack messages discussing the deletions — a ruling that, if the messages show intent to obscure infringing sourcing, could expose OpenAI to enhanced, willfulness-based damages later in the case.

Where it stands

Fact discovery closed on February 27, 2026. The court has since pushed the summary judgment schedule, with reply briefs now due November 6, 2026 rather than the original October date — a sign of how much material the parties are still working through. No ruling on the merits of infringement, fair use, or class certification has issued.

Why it matters for training-data licensing

This is the largest author class action in the space, and the deleted-dataset fight is the part that should worry any data buyer: a decision to delete source records, made years before litigation, is now itself potential evidence of consciousness of wrongdoing. For companies assembling training corpora today, the lesson is blunt — keep the receipts. Provenance records that look irrelevant at collection time can become the difference between an ordinary fair-use defense and a spoliation problem in front of a jury.

What to watch

  • Summary judgment briefing, with replies due November 6, 2026.
  • Whether the Slack messages about the deleted datasets support findings of willfulness or spoliation.
  • Any class certification motion following the close of fact discovery.
  • Whether the output-infringement theory becomes a template for other author suits inside the MDL.

Sources

copyrightbooksclass actionfair usediscoverytraining data

See something wrong? Send a correction.

Jaeden Schafer

Jaeden Schafer

Jaeden Schafer is the founder of fiund and host of the AI Chat podcast. He covers the training-data market and the lawsuits shaping it.

Want data that clears this in diligence?

Whether you're building a model or sitting on an archive, the first conversation is short and specific.

Send a brief