Rights & provenance

Training-data copyright lawsuits: the landscape

Last updated 2026-07-21

⚠ Legal specifics move monthly — re-verify against primary sources before relying on this page.

A running map of the major cases testing whether training AI on copyrighted material without a licence is infringement or fair use.

Why it matters to a buyer

The outcomes directly price the risk of unlicensed data. Buyers need an accurate read, not headlines.

Why it matters to a data owner

The litigation is why licensed supply is worth more — it is the alternative to legal exposure.

Current legal status

The picture is nuanced and case-specific. Thomson Reuters v. Ross (Feb 2025) was the first US ruling to reject a fair-use defense for training data — but it is a non-generative case, so its reach to generative LLMs should not be overstated. Getty v. Stability in the UK High Court (Nov 2025) was largely a loss for Getty, which had withdrawn its primary training claims for lack of UK jurisdiction and lost the secondary-infringement claim, retaining only a narrow trademark win — it is not a rights-holder win on training. Other cases (NYT v. OpenAI, UMG v. Anthropic) remain live. [Verify every outcome against primary sources before relying on it.]

What fiund does about it

fiund exists so buyers never have to bet a model on how these cases resolve.

Sources

← All rights & provenance guides

More on rights & provenance

Want data that clears this in diligence?

Whether you're building a model or sitting on an archive, the first conversation is short and specific.

Send a brief