Lawsuit tracker
S.D.N.Y. · ActivePublishers v. Cohere
A judge let publishers' full slate of claims against Cohere survive the pleadings stage, including a theory that AI-generated summaries can infringe even without verbatim copying. The case is now in discovery.
Key facts
- Fourteen publishers — including Condé Nast, The Atlantic, Forbes, The Guardian, Politico, Vox Media, Business Insider, the Los Angeles Times and Advance Local Media — sued Cohere in the S.D.N.Y. on February 13, 2025.
- The complaint, backed by the News/Media Alliance, identifies more than 4,000 works and alleges Cohere's Command models were trained on and can output near-verbatim or substitutive summaries of the publishers' articles, sometimes from behind paywalls.
- It also alleges Command hallucinates articles and falsely attributes them to publisher brands, pleaded as Lanham Act trademark claims.
- On November 13, 2025, Judge Colleen McMahon denied Cohere's partial motion to dismiss in full, holding that substitutive summaries can plausibly infringe even without verbatim copying.
- Cohere answered the complaint in December 2025, and the case is now in discovery with no trial date set.
The allegations
Filed February 13, 2025, this is the first broad industry-coalition suit against an enterprise large-language-model vendor rather than a consumer-facing chatbot company. The fourteen publishers allege Cohere trained its Command family of models on their journalism without a license and that Command can reproduce substantial portions of specific articles — sometimes near-verbatim, sometimes as detailed summaries the publishers say substitute for the original in the market, and sometimes drawn from paywalled content. The complaint separately alleges Command hallucinates entirely fabricated articles and misattributes them to the plaintiffs' mastheads, which the publishers plead as false designation of origin and trademark dilution under the Lanham Act. They seek statutory damages of up to $150,000 per infringed work, covering more than 4,000 identified pieces, plus an injunction.
The paywall allegations matter on their own: several of the identified excerpts allegedly reproduced content originally locked behind a subscription wall, which the publishers frame as a distinct market-harm theory from ordinary reproduction of freely available web content — a model that lets users bypass the exact access controls a publisher relies on for subscription revenue does a different kind of economic damage than one that merely echoes public web pages.
Cohere's business model is part of what makes this case distinct: it licenses Command chiefly to enterprise customers who build their own products on top of the model, rather than running a mass-market consumer chatbot itself. That distribution structure means litigation risk here flows downstream — any company that fine-tuned or deployed Command inside its own product inherits exposure tied to Cohere's training data, which is exactly the kind of risk that indemnification clauses in enterprise AI contracts are increasingly written to address.
The legal question
The case turns on a question most AI copyright suits haven't squarely presented: can an AI-generated summary infringe copyright even when it isn't a verbatim copy? Cohere's defense leaned on the idea that summarizing is different in kind from copying. The publishers countered that a summary detailed enough to substitute for the original article in the market causes the same competitive harm copyright is meant to prevent, regardless of exact wording. A second question rides alongside it: does an AI system inventing fake articles and slapping a real publisher's name on them create trademark liability distinct from the copyright claims?
The ruling and its limits
On November 13, 2025, Judge McMahon denied Cohere's partial motion to dismiss in its entirety. She held the publishers plausibly stated a claim that substitutive AI summaries can infringe even without verbatim copying, and let both the secondary-liability theories and the hallucination-based trademark claims proceed. It was a clean sweep for the publishers at the pleadings stage.
That ruling establishes only that the claims are legally viable enough to proceed — it is not a finding that Cohere actually infringed any of the 4,000-plus identified works, and no court has yet examined the outputs at issue in detail or ruled on Cohere's fair-use defense. Cohere answered the complaint in December 2025, and the case moved into discovery on that basis.
Where it stands
As of July 2026, the case is active and in discovery, with no trial date set and no ruling on the merits of infringement or fair use.
Why it matters for training-data licensing
This was the first ruling to squarely reject "we summarize, we don't copy" as a safe harbor: a summary detailed enough to replace the article it describes can infringe on its own terms. That matters well beyond Cohere, because summarization is a default behavior of nearly every deployed language model, not an edge case.
It's also a reminder that enterprise model vendors — the companies licensing models into other products, rather than running consumer chatbots — are squarely in scope for the same training-data claims. And the hallucination theory adds a layer buyers should track separately from copyright: a model that fabricates content and attributes it to a real publisher creates trademark-shaped risk that survives even if the underlying training-data question is eventually resolved in the vendor's favor. Anyone licensing or fine-tuning on news content should treat both risks as live, not just the copyright one.
What to watch
- Discovery into the more than 4,000 specific works identified in the complaint.
- Whether the substitutive-summary theory survives summary judgment, not just the pleadings.
- Coordination with other publisher suits, including the NYT MDL and the newly filed Google Gemini case.
- Whether other enterprise LLM vendors face similar hallucination-based trademark claims.
Sources
See something wrong? Send a correction.
More case files
Want data that clears this in diligence?
Whether you're building a model or sitting on an archive, the first conversation is short and specific.
Send a brief