Rights & provenance
Training-data copyright lawsuits: the landscape
Last updated 2026-07-21
A running map of the major cases testing whether training AI on copyrighted material without a licence is infringement or fair use.
Why it matters to a buyer
The outcomes directly price the risk of unlicensed data. Buyers need an accurate read, not headlines.
Why it matters to a data owner
The litigation is why licensed supply is worth more — it is the alternative to legal exposure.
Current legal status
The picture is nuanced and case-specific. Thomson Reuters v. Ross (Feb 2025) was the first US ruling to reject a fair-use defense for training data — but it is a non-generative case, so its reach to generative LLMs should not be overstated. Getty v. Stability in the UK High Court (Nov 2025) was largely a loss for Getty, which had withdrawn its primary training claims for lack of UK jurisdiction and lost the secondary-infringement claim, retaining only a narrow trademark win — it is not a rights-holder win on training. Other cases (NYT v. OpenAI, UMG v. Anthropic) remain live. [Verify every outcome against primary sources before relying on it.]
What fiund does about it
fiund exists so buyers never have to bet a model on how these cases resolve.
Sources
More on rights & provenance
Want data that clears this in diligence?
Whether you're building a model or sitting on an archive, the first conversation is short and specific.
Send a brief