Lawsuit tracker
Britannica v. Perplexity
Encyclopaedia Britannica, Inc. and Merriam-Webster, Inc. v. Perplexity AI, Inc., No. 1:25-cv-07546 (S.D.N.Y.)
| Plaintiffs | Encyclopaedia Britannica, Inc.; Merriam-Webster, Inc. |
|---|---|
| Defendants | Perplexity AI, Inc. |
| Court | U.S. District Court, Southern District of New York (Judge Jennifer L. Rochon) |
| Filed | 2025-09-10 |
| Status | Active |
| Content type | reference works |
| Last updated | 2026-07-21 |
The claims
Copyright infringement on both the input side (crawling and scraping) and the output side (verbatim or near-verbatim reproduction in answers), plus trademark infringement over attributing AI-garbled content to the Britannica and Merriam-Webster names
What has happened
Britannica and its Merriam-Webster unit sued Perplexity in September 2025 over the “answer engine.” The complaint attacks two stages: PerplexityBot scraping their sites to feed its index, and the retrieval-augmented outputs that reproduce encyclopedia entries and dictionary definitions, sometimes verbatim; the complaint points to Perplexity returning Merriam-Webster’s definition of “plagiarize.” The trademark counts add a different injury: Perplexity allegedly attaches the Britannica and Merriam-Webster names to inaccurate, AI-generated text, trading on and damaging the brands. In March 2026 the plaintiffs amended their complaint, and Perplexity moved to dismiss the output-based direct infringement claim, arguing users, not Perplexity, supply the prompts and thus any volitional conduct. The plaintiffs countered that Perplexity cannot shift liability for its outputs onto its users. The motion is pending.
Key developments
- 2025-09-10 — Britannica and Merriam-Webster sue Perplexity in the Southern District of New York for copyright and trademark infringement over its answer engine.
- 2026-03-17 — Plaintiffs file an amended complaint; Perplexity simultaneously moves to dismiss the output-side direct infringement claim on volitional-conduct grounds. Briefing followed; decision pending.
Why it matters for training data
The battleground here is retrieval, not training. If reproducing reference entries in AI answers is direct infringement, answer engines need display and retrieval licenses even where model training might separately qualify as fair use, which splits “training rights” from “output rights” in any data deal. The trademark theory is the second lever: when an AI product garbles branded content and still cites the brand, the licensor has a claim that has nothing to do with copying. Reference publishers whose value is accuracy have the most to gain from that theory, and buyers should price both rights separately.
Sources
- Docket, No. 1:25-cv-07546 (CourtListener)
- Britannica corporate announcement of the lawsuit
- Complaint (Susman Godfrey)
- MLex: Perplexity urges dismissal of output infringement claim
Deeper analysis
Want data that clears this in diligence?
Whether you're building a model or sitting on an archive, the first conversation is short and specific.
Send a brief