Rights & provenance

What diligence buyers actually run on training data

Last updated 2026-07-22

A practical walk-through of the questions a foundation-model legal team asks before licensing a dataset.

Why it matters to a buyer

Knowing the checklist lets you evaluate any vendor — including fiund — on the axis that matters.

Why it matters to a data owner

Understanding buyer diligence helps owners prepare documentation that clears it.

Current legal status

Practice-based, not statutory.

The five questions, in the order they get asked

Diligence on a training dataset tends to run the same arc regardless of the buyer. First, rights: who owns this material, and does the licence on offer actually grant model training in words a lawyer can point to? Second, people: does anyone identifiable appear or speak, and if so, where is their consent? Third, acquisition: how was the material collected — recorded by the owner, transferred, scraped, bought — and can each step be shown? Fourth, regulatory surface: where are the people in the data located, because EU speakers raise GDPR questions and Illinois speakers raise biometric-consent questions that copyright paperwork does not answer. Fifth, documentation: is there a provenance record that ties all of the above to the actual files, or only assertions?

The order matters because each question gates the next. A perfect consent file is worthless if the licensor never had the rights to grant; a clean licence is fragile if the material underneath it was scraped. Courts have made the acquisition question decisive — the Bartz v. Anthropic split between training on lawfully acquired books and keeping a pirated library, followed by a $1.5 billion settlement, is the reason "where did this come from" now leads the list.

A worked example: the request list for a speech corpus

A lab evaluating a licensed speech dataset typically asks the supplier for a document set like this: the licence chain — the capture or production agreement, any transfer or licence at each change of hands, and the licence being offered to the buyer, so the grant can be traced end to end; sample consent forms plus confirmation of coverage for every identifiable speaker, with the AI use named; a description of the collection methodology — who recorded what, where, on what terms; speaker locations, to scope GDPR and biometric-statute exposure; provenance records with file hashes and timestamps; and the warranty schedule the supplier is prepared to stand behind.

And the red flags that end deals: "it was publicly available" offered as a rights answer; consent language that predates AI or never names training; a licensor who cannot explain how the material was acquired; chain-of-title gaps around exactly the material that matters; and a supplier who cannot say where its speakers are located.

Why the checklist got harder in 2025-26

Two shifts turned this from a lawyer’s preference into a pipeline requirement. The litigation wave priced sloppy acquisition — Bartz ended in a $1.5 billion settlement after the fair-use answer split on provenance. And the EU AI Act made documentation regulatory: since August 2, 2025, providers of general-purpose models placed on the EU market must publish a summary of training content and maintain a copyright policy honoring rights reservations, with direct Commission enforcement from August 2, 2026. A buyer with EU exposure now collects in diligence exactly what it must later describe in public — which is why documentation demands flow down to every supplier.

The document set that clears diligence

What suppliers are commonly asked to produce, in one list:

  • A signed licence with an explicit AI-training grant and stated scope.
  • The full licence chain: capture agreement, any transfers, and the licence to the buyer.
  • Consent artifacts for every identifiable person, naming AI training as a use.
  • Collection methodology: how, where, and by whom the material was gathered.
  • Speaker and subject locations, for GDPR and biometric-statute analysis.
  • Provenance records: source identity, timestamps, and file hashes.
  • A warranty schedule covering ownership, non-infringement, and lawful collection.

What fiund does about it

fiund is built to pass this checklist: signed licence, explicit training rights, separate consent, no scraping, provenance on request.

← All rights & provenance guides

Frequently asked questions

What single document matters most in diligence?

The licence grant. Review starts with whether model training is named as a permitted use and what scope was granted — everything else exists to verify that the party granting it could, and that the people in the data consented.

Do buyers verify provenance, or just take warranties?

Documentation first, warranties second. A provenance record a legal team can verify beats a promise the buyer would have to sue on — warranties allocate the residual risk, they do not replace the record.

What are the most common ways a dataset fails diligence?

No explicit training grant; consent that is missing or never names AI; acquisition nobody can explain; gaps in the chain of title; and no answer to where the speakers are located.

Does diligence change if the model will be available in the EU?

Yes. Providers of general-purpose models on the EU market must publish a training-content summary and maintain a copyright policy under Article 53 of the EU AI Act, so buyers collect supplier documentation they will later have to stand behind publicly.

More on rights & provenance

Want data that clears this in diligence?

Whether you're building a model or sitting on an archive, the first conversation is short and specific.

Send a brief