Lawsuit tracker

S.D.N.Y. · Active

Reddit v. Perplexity

Reddit's second scraping suit skips past Perplexity and goes straight for the proxy networks that fed it. If a downstream buyer of scraped data can be liable under the DMCA, "we bought it from a vendor" stops working as a defense.

Key facts

  1. Reddit, Inc. v. SerpApi, LLC, Oxylabs UAB, AWMProxy and Perplexity AI, Inc., No. 1:25-cv-08736, filed October 22, 2025 in the Southern District of New York before Judge Paul A. Engelmayer.
  2. The complaint alleges Digital Millennium Copyright Act anti-circumvention violations plus state-law unjust enrichment, and names three scraping intermediaries alongside Perplexity.
  3. Reddit alleges SerpApi, Oxylabs and AWMProxy harvested its content at what the complaint calls industrial scale by scraping Google search results and masking their identities to evade Reddit's technical blocks, then sold the data on to Perplexity.
  4. Perplexity denies training on Reddit content and says it summarizes and cites public discussions; it and SerpApi moved to dismiss the amended complaint in March 2026.
  5. Reddit filed its consolidated opposition on April 17, 2026. Oral argument on the motions to dismiss is set for June 30, 2026, with a decision pending as of July 2026.
PartiesReddit, Inc. v. Perplexity AI, Inc.; SerpApi, LLC; Oxylabs UAB; AWMProxy
CourtU.S. District Court, Southern District of New York (Judge Paul A. Engelmayer)
Docket1:25-cv-08736
Filed2025-10-22
Content typeplatform data
StatusAs of July 2026, Perplexity’s and SerpApi’s motions to dismiss the amended complaint are pending before Judge Engelmayer, with argument set for June 30, 2026.

The second Reddit suit goes after the supply chain

Reddit already has one AI training-data suit running, against Anthropic, over contract terms and API scraping. This second suit, filed in October 2025, is a different animal. It doesn't stop at the AI company answering user queries with Reddit content. It reaches back through the pipeline to the companies that did the scraping in the first place.

Named alongside Perplexity are SerpApi, a search-results API reseller, Oxylabs, a proxy-network operator, and AWMProxy, another proxy service. None of them are AI labs. All three, Reddit alleges, exist to get around access controls other companies put up — and Reddit put up plenty after it began charging for API access in 2023.

What Reddit says happened

Reddit's theory: locked out of direct access, the intermediaries scraped Reddit content as it appears in Google search results instead, at what the complaint calls industrial scale, while disguising their identity to dodge Reddit's blocking measures. That is the circumvention Reddit says violates the DMCA's anti-circumvention provision, separate from ordinary copyright infringement.

Perplexity, Reddit alleges, bought the resulting dataset from at least one of the intermediaries and now surfaces Reddit content in its answer engine — summaries, citations and excerpts pulled from threads Reddit never licensed to Perplexity directly. Perplexity disputes that framing and says it summarizes and cites public discussions the way any search-adjacent product would.

Perplexity's downstream-buyer defense

Perplexity and SerpApi moved to dismiss in March 2026. Perplexity's core argument: it sits downstream of any circumvention that may have occurred, the DMCA's anti-circumvention provision doesn't carry secondary liability the way ordinary copyright infringement does, and simply buying data that someone else scraped is not itself a violation of the statute.

Reddit's April 17, 2026 consolidated opposition pushes back on all three points, arguing that knowingly buying circumvention-derived data and monetizing it is not meaningfully different from doing the circumventing yourself. That is now the live question in front of Judge Engelmayer.

Provenance and the 'we bought it from a vendor' defense

This is the case to watch for anyone buying, not scraping, training data. Most training-data suits target the company that did the collecting. This one asks whether the DMCA reaches a company two or three steps removed — one that never touched Reddit's servers, only a vendor's output file.

If that liability theory survives the motion to dismiss, a vendor's assurance that data was "sourced legally" stops being enough. Buyers will need actual chain-of-title documentation: where the data came from, whether the seller had the right to collect it, and whether any technical protection measure was bypassed upstream. Contract warranties about lawful sourcing move from boilerplate to load-bearing.

What's next

Argument on the motions to dismiss is set for June 30, 2026. A ruling for Perplexity would push the fight onto the intermediaries and strengthen the "downstream buyer" defense industry-wide. A ruling for Reddit would put every company that licenses aggregated web data on notice that its supplier's methods are now its own legal exposure.

What to watch

  • The outcome of the June 30, 2026 oral argument and whether Judge Engelmayer lets the case proceed against Perplexity as a downstream buyer.
  • Whether the DMCA's anti-circumvention provision is read to reach purchasers of already-scraped data, not just the parties who did the scraping.
  • Whether SerpApi, Oxylabs or AWMProxy settle separately, given they sit closer to the alleged circumvention than Perplexity does.
  • Any read-across to Reddit's parallel, contract-based suit against Anthropic over API terms.

Sources

RedditPerplexityDMCAData scrapingData provenance

See something wrong? Send a correction.

Jaeden Schafer

Jaeden Schafer

Jaeden Schafer is the founder of fiund and host of the AI Chat podcast. He covers the training-data market and the lawsuits shaping it.

Want data that clears this in diligence?

Whether you're building a model or sitting on an archive, the first conversation is short and specific.

Send a brief