Lawsuit tracker

S.D.N.Y. · Active

Dow Jones v. Perplexity

News Corp’s Wall Street Journal and New York Post sued Perplexity over its answer engine, pairing copyright claims with trademark theories over hallucinated citations. Perplexity’s bid to escape New York failed, and the case is now deep in discovery.

Key facts

  1. Filed October 21, 2024 in the Southern District of New York as Dow Jones & Co., Inc. and NYP Holdings, Inc. v. Perplexity AI, Inc., No. 1:24-cv-07984, before Judge Katherine Polk Failla.
  2. Claims cover copyright infringement — both copying articles into Perplexity’s index and outputs that reproduce or closely paraphrase them — plus trademark dilution and false-designation claims over answers, including hallucinated content, attributed to the Journal and the Post.
  3. Perplexity moved on February 18, 2025 to dismiss for improper venue or transfer the case to California; Judge Failla denied the motion in full on August 21, 2025.
  4. The court held Perplexity’s consumer user-agreement forum clause does not bind publishers who never used the product as customers.
  5. On March 20, 2026, Failla ordered Perplexity to search its founders’ personal email accounts used for company business, covering January 2022 through February 2026, with production due July 15, 2026.
PartiesDow Jones (publisher of The Wall Street Journal) and NYP Holdings (New York Post), both News Corp companies v. Perplexity AI
CourtU.S. District Court, Southern District of New York (Judge Katherine Polk Failla)
Docket1:24-cv-07984
Filed2024-10-21
Content typejournalism
StatusAs of July 2026, in active discovery after Perplexity lost its bid to dismiss or move the case to California. Fact discovery was scheduled to close July 20, 2026, with expert discovery running to October and a pretrial conference set for December 14, 2026.

What’s alleged

Both plaintiffs are News Corp properties — Dow Jones publishes the Wall Street Journal and Barron’s, and NYP Holdings publishes the New York Post — giving the suit the backing of a major publisher pursuing an AI company over its live product rather than a one-time training run. Filed in October 2024, it was among the earliest suits to target an AI answer engine specifically, rather than a foundation-model developer’s training pipeline in isolation.

Dow Jones and the New York Post allege Perplexity’s retrieval-augmented answer engine copies their articles wholesale into a search index and then serves outputs that substitute for reading the originals — diverting traffic and subscription revenue. The publishers seek statutory damages of up to $150,000 per infringed work.

Layered on top is a trademark theory distinctive to this case: false-designation and dilution claims over answers, including allegedly hallucinated ones, that Perplexity presents under the Journal’s or the Post’s name. The publishers argue that misattributing fabricated content to their mastheads harms their brands independent of any copyright harm.

The legal question: is retrieval a shortcut around training-data copyright fights?

Perplexity’s product doesn’t just train on articles once and generate from memory — it crawls and indexes content on an ongoing basis, then answers queries with excerpts and summaries in something close to real time. That raises a question distinct from the training-only cases: whether a retrieval-augmented product’s continuous copying into an index gets analyzed the same way as embedding works into model weights during a training run, or differently, since the copying is more direct and more current.

The publishers frame their case around substitution — a user who gets the gist of an article from Perplexity’s answer has less reason to visit the original — which goes directly to market harm, the factor most likely to matter if the case ever reaches a fair-use defense. That distinction is not academic: courts assessing copyright claims against generative tools typically weigh whether an output substitutes for the original in the market, and a product whose entire value proposition is delivering summarized news without a click-through is a closer fit for that theory than a model that occasionally paraphrases a training example from years earlier. No fair-use ruling has been reached; the case remains in discovery.

The ruling and its limits: no exit from New York

Perplexity’s February 2025 motion sought dismissal for lack of personal jurisdiction and improper venue, or transfer to the Northern District of California, where the company is headquartered. Judge Failla denied the motion in full in August 2025, pointing to Perplexity’s own New York office, staff and marketing activity as sufficient contacts with the forum.

The ruling also disposed of a venue argument built on Perplexity’s consumer terms of service, which route disputes through the company’s preferred forum. Failla held those terms don’t bind rights holders who never accepted them by using the product — a limit that matters well beyond this case, since AI companies increasingly rely on EULA forum clauses to control where they get sued.

Perplexity’s consumer terms of service, which route disputes through the company’s preferred forum, do not bind rights holders who never agreed to use the product.— Order Denying Motion to Dismiss or Transfer, No. 1:24-cv-07984 (S.D.N.Y. Aug. 21, 2025)

Where it stands: discovery, including founders’ personal email

The case is in active fact discovery. At a March 20, 2026 conference, Judge Failla ordered Perplexity to search its founders’ personal email accounts — where company business was allegedly conducted — for the period from January 2022 through February 2026, with production due July 15, 2026. Orders reaching into personal accounts are unusual and signal the court’s willingness to dig into internal decisions about scraping and indexing practices.

No trial date has been reported. The dispute remains squarely in the fact-gathering phase, with the founders’ email production the most consequential near-term milestone.

Why it matters for training-data buyers

Retrieval is not a loophole: grounded, cite-and-summarize products face the same copyright exposure as training pipelines, with trademark exposure stacked on top whenever outputs misattribute content, hallucinated or not. Buyers evaluating a RAG-based product should treat content licensing for the live index with the same seriousness as licensing for model training, not as a separate, lower-stakes question.

The venue ruling adds a chain-of-title-adjacent lesson: a vendor’s consumer EULA cannot be used to drag a rights holder who never used the product into a friendlier forum. For publishers and other content owners negotiating with AI companies, that means litigation leverage sits with the rights holder’s home jurisdiction, not the vendor’s terms of service.

For companies buying a RAG product off the shelf, the practical procurement question is narrower than it looks: don’t just ask whether a vendor licensed the data behind its model. Ask whether the live retrieval corpus it crawls and re-crawls to answer queries carries its own licensing coverage, since this case treats indexing as a potentially infringing act in its own right, separate from any one-time training run.

What to watch

  • Perplexity’s July 15, 2026 production of founders’ personal email records covering January 2022 through February 2026.
  • Whether the case reaches a fair-use ruling on retrieval-augmented outputs, a question distinct from training-only cases.
  • Any additional publisher plaintiffs joining or filing parallel suits against answer-engine products.
  • A trial date, which has not yet been set as of July 2026.

Sources

journalismRAGtrademarkS.D.N.Y.Perplexity

See something wrong? Send a correction.

Jaeden Schafer

Jaeden Schafer

Jaeden Schafer is the founder of fiund and host of the AI Chat podcast. He covers the training-data market and the lawsuits shaping it.

Want data that clears this in diligence?

Whether you're building a model or sitting on an archive, the first conversation is short and specific.

Send a brief