Lawsuit tracker

S.D.N.Y. · Active

Britannica v. Perplexity

Britannica and Merriam-Webster say Perplexity’s answer engine does not just scrape their reference works — it reproduces them at query time and attaches the brand name to inaccurate, AI-written text.

Key facts

  1. Filed September 10, 2025 in S.D.N.Y. by Encyclopaedia Britannica and its Merriam-Webster unit against Perplexity AI.
  2. Claims copyright infringement on both the input side (crawling) and output side (near-verbatim reproduction of entries and definitions), plus trademark infringement.
  3. The complaint cites Perplexity returning Merriam-Webster’s definition of “plagiarize” as an example of verbatim output.
  4. Plaintiffs amended their complaint March 17, 2026; Perplexity simultaneously moved to dismiss the output-based direct-infringement claim on volitional-conduct grounds.
  5. The motion to dismiss is pending before Judge Jennifer Rochon.
PartiesEncyclopaedia Britannica, Inc.; Merriam-Webster, Inc. v. Perplexity AI, Inc.
CourtU.S. District Court, Southern District of New York (Judge Jennifer L. Rochon)
Docket1:25-cv-07546
Filed2025-09-10
Content typereference works
StatusAs of July 2026, Perplexity’s motion to dismiss the output-side direct infringement claim (filed March 17, 2026, alongside an amended complaint) awaits decision; the case is otherwise proceeding.

Two theories: scraping and answers

Britannica and Merriam-Webster split their case into two distinct acts of alleged infringement. The first is familiar: PerplexityBot crawling their sites to build a retrieval index. The second is what sets this case apart from most training-data suits — it targets what Perplexity’s “answer engine” does at query time, arguing its retrieval-augmented answers reproduce encyclopedia entries and dictionary definitions closely enough, sometimes verbatim, to infringe independently of anything that happened during training.

The volitional-conduct defense

In March 2026, alongside an amended complaint from the plaintiffs, Perplexity moved to dismiss the output-based claim. Its argument leans on a doctrine called volitional conduct: direct copyright infringement generally requires a defendant to take an active, causal role in the copying, not simply provide a tool that a user directs. Perplexity says its users supply the prompts that trigger any given answer, so any copying is triggered by user action rather than the company’s own conduct — the same argument search engines and cloud-storage providers have raised for reproduction claims in the past. Britannica and Merriam-Webster counter that Perplexity built and controls the retrieval and generation pipeline, and cannot offload responsibility for its own product’s outputs onto the people who type into it.

Why trademark, not just copyright

The trademark claims add an injury copyright law does not reach: reputational harm from misattribution. The complaint alleges Perplexity’s answers sometimes attach the Britannica or Merriam-Webster name to inaccurate or garbled AI-generated text, which — regardless of whether the underlying words are copied — can mislead users about the source and quality of the information and damage the brands’ reputation for accuracy.

Where it stands

The plaintiffs’ amended complaint and Perplexity’s motion to dismiss the output claim were filed together on March 17, 2026, and briefing has followed. Judge Jennifer Rochon has not yet ruled. Because the motion targets only the output-based direct-infringement count, the crawling-based claims and the trademark claims are not before the court on this motion.

Why it matters for training-data licensing

If courts accept that retrieval-based outputs can independently infringe, licensing a model’s training data will not be enough to clear an answer engine — display and retrieval rights become a separate license to negotiate, distinct from training rights. That splits what used to be one deal into two, and it particularly benefits reference and news publishers whose value is precision: a garbled paraphrase attributed to a trusted name is its own kind of harm, the one the trademark claim is built to capture, independent of how any copyright question comes out.

What to watch

  • Judge Rochon’s ruling on Perplexity’s motion to dismiss the output-infringement claim.
  • Whether other answer-engine products face parallel volitional-conduct arguments.
  • Discovery into how closely Perplexity’s retrieval pipeline mirrors source text before generation.
  • Any separate licensing deals reference publishers strike with AI search products while the case proceeds.

Sources

copyrighttrademarkretrievalreference publishingfair use

See something wrong? Send a correction.

Jaeden Schafer

Jaeden Schafer

Jaeden Schafer is the founder of fiund and host of the AI Chat podcast. He covers the training-data market and the lawsuits shaping it.

Want data that clears this in diligence?

Whether you're building a model or sitting on an archive, the first conversation is short and specific.

Send a brief