Rights & provenance

Is scraped data fair use for AI training?

Last updated 2026-07-21

Whether scraping copyrighted material to train a model is fair use is the central legal question of the space — and it is unsettled.

Why it matters to a buyer

Betting a shipped model on an unsettled fair-use defense is a business risk, not just a legal one.

Why it matters to a data owner

The uncertainty is precisely why licensed material has a market.

Current legal status

There is no blanket ruling that training on scraped or otherwise unlicensed data is or is not fair use, and the 2025 district-court decisions split on their facts: Thomson Reuters v. Ross rejected fair use for a non-generative tool that substituted directly for the source; Bartz v. Anthropic found training on lawfully acquired books fair use while holding a pirated library was not; Kadrey v. Meta granted Meta summary judgment on the record before it, while warning that stronger market-dilution evidence could change the outcome in another case. Outcomes turn on how the data was acquired, what the system does, and market harm. The first appellate answer is pending in the Third Circuit after argument on June 11, 2026.

What fiund does about it

Licensing removes the question entirely.

Sources

← All rights & provenance guides

More on rights & provenance

Want data that clears this in diligence?

Whether you're building a model or sitting on an archive, the first conversation is short and specific.

Send a brief