Resources

What to check before licensing a dataset

A pre-purchase diligence pass for AI teams licensing audio, video, or UGC. Consent records are the hard stop.

Published 2026-07-22 · 7 min read

Key takeaways

  1. Quality and rights are different questions. Answer rights first.
  2. A signed AI-training licence is the floor, not stock terms or a checkbox.
  3. Consent for identifiable people is the one check with no workaround.
  4. Match the sourcing story to the documents. Vague sourcing means vague title.
  5. Read the indemnity limits, not just the fact that one exists.

A dataset can sound clean, arrive in the right format, and still be unsafe to train on. The sample tells you about quality. It tells you almost nothing about rights. Before you license anything, separate the two questions and answer the rights question first.

This is a pre-purchase pass. Run it before money moves, and before the data touches your pipeline. Most of it is document review, not listening tests. If a supplier cannot produce the documents, that is your answer.

Start with the grant, not the sample

Every asset you train on needs a signed licence that permits AI training. Not a stock licence. Not a terms-of-service checkbox. A written grant that names model training as a permitted use.

Read the grant for scope. Does it cover training, fine-tuning, evaluation, and derivative models. Does it cover the modalities you actually need. Does it survive after the deal ends, or must you delete trained artifacts.

On fiund, a signed AI-training licence sits on every asset by default. Your job in diligence is to read it, not to assume it.

Consent records are the hard stop

If a person is identifiable in the data, a licence from the owner is not enough. You also need that person’s consent for their voice or likeness to be used in training.

A studio can own a recording and still lack the right to license the speaker’s voice for AI. Ownership of the file and consent of the person are two different things.

Ask for the consent records. Ask how identifiable people were informed, what they agreed to, and whether they can withdraw. If the supplier cannot show consent for identifiable subjects, stop. This is the one check with no workaround. On fiund, separate voice and likeness consent is attached where people are identifiable, kept alongside the licence.

Chain of title and sourcing method

Chain of title is the unbroken line from the person who created the content to the party selling it to you. If a link is missing, the grant may not be theirs to give.

Ask how the data was sourced. Commissioned and paid. Contributed by consenting creators. Drawn from an archive with clear ownership. Each of these can be clean. Scraped from the open web, relabelled, and resold is not.

Match the story to the documents. A supplier who describes commissioned capture should be able to show contributor agreements. A vague answer here usually means a vague chain of title.

Exclusivity, deliverable spec, and indemnity

Exclusivity changes what you are buying. Non-exclusive is the common default, and it is often fine. But know it. If a competitor can license the same set, decide whether that matters for your use before you sign, not after.

Pin down the deliverable spec in writing. Hours or clips. Formats and codecs. Sample rate. Metadata fields. Label schema. Known gaps. A dataset that is right on rights and wrong on spec still fails.

Read the indemnity. A supplier confident in their rights will stand behind them. Look for indemnification against third-party IP and consent claims, and check the limits. An indemnity capped near zero is a signal, not a protection.

What a clean answer looks like

Good suppliers expect these questions. They hand over the licence, the consent records, and a provenance trail without friction.

On fiund, owners keep ownership and approve buyers, licences are non-exclusive by default, and provenance is available during diligence. The point of the pass is not suspicion. It is proof.

Watch outIf a supplier cannot produce consent records for identifiable people, do not proceed. Ownership of a recording is not the same as consent to train on a person’s voice or face.

Pre-purchase diligence checklist

  • Signed AI-training grant that names model training as a permitted use
  • Consent records for every identifiable person, covering voice and likeness
  • An unbroken chain of title from creator to seller
  • A sourcing method the supplier can describe and document
  • Exclusivity terms stated in writing, with non-exclusive as a common default
  • A deliverable spec: volume, formats, metadata, label schema, known gaps
  • Indemnification against third-party IP and consent claims, with readable limits

← All resources

Frequently asked questions

Is a stock or content licence enough to train a model?

Usually not. Most stock and content licences predate AI training and do not grant it. You need a licence that names model training as a permitted use. If it is silent, treat it as not granted.

The people in the clips are not famous. Do I still need consent?

Yes. Consent is about identifiability, not fame. If a face or voice can be recognised, you need that person’s consent for training use, separate from the owner’s licence.

What if I only need the data for evaluation, not training?

Say so in the grant. Evaluation, fine-tuning, and training are different uses. A licence that covers one may not cover the others. Scope the grant to what you will actually do.

Related resources

Want data that clears this in diligence?

Whether you're building a model or sitting on an archive, the first conversation is short and specific.

Send a brief