Resources

Can I sell data that has other people in it?

You own the recording. Other people are in it. Here is when that is fine, when you need a release, and when to leave a clip out.

Published 2026-07-22 · 6 min read

The landscape here moves quickly — re-verify against the cited sources for the current status before you rely on it.

Key takeaways

  1. Owning a recording is not the same as holding the rights to every voice in it.
  2. Identifiable people carry personal rights — publicity in the US, data protection in the UK and EU — that need their consent.
  3. Background crowd wash is lower risk; a clear, quotable bystander is not.
  4. When you cannot clear a voice, exclude that clip rather than risk the set.
  5. Per-asset consent lets you license the clean material and leave the rest out.

You own the recording. You control the file. But other people are in it — a co-host, a guest, a caller, someone in the background. So can you license it for AI training?

Often yes. Sometimes only after you clear the other voices. The test is not who owns the file. It is whether the people in it are identifiable, and whether they agreed to this use.

One file, two kinds of rights

A recording carries two separate rights. There is the copyright in the recording itself, which you may own. And there are the personal rights of the people captured in it — their voice, their likeness, their personal data.

Owning the copyright lets you license the recording. It does not, on its own, let you license someone else’s voice or face for training a model. Those personal rights stay with the person. In the United States they sit under the right of publicity and state privacy law. In the United Kingdom and European Union they sit under data-protection law.

So the question splits in two. Do you have the right to license the content? And do the identifiable people in it consent to that use? You need both.

Who counts as identifiable

Identifiable does not only mean famous. It means a person can be recognised. A named guest is identifiable. A caller who gives a first name and a town is identifiable. A voice that regular listeners would know is identifiable.

Under data-protection rules, a voice recording is personal data when it relates to a person who can be identified, directly or indirectly. That is a low bar. Most speech in a podcast or interview clears it.

If a person is identifiable, their consent matters. If they are genuinely not — a distant, indistinct murmur in a crowd — the risk is lower. But do not guess. When in doubt, treat the person as identifiable.

Bystanders and background voices

Field recordings and street audio raise the hardest version of this. You set out to record ambience. You also caught a bystander’s conversation.

A crowd wash where no single person can be picked out is one thing. A clear, quotable exchange from one passer-by is another. The second is closer to captured speech from an identifiable person, even if you never learned their name.

The practical fix is capture discipline. Note where consent was given. Flag clips where it was not. Keep the two apart so a non-consenting voice never ends up in a licensed set.

When you need a release

You need a release, or documented consent, when an identifiable person’s voice or likeness is in material you want to license for AI training. A guest, a co-host, an interviewee, a named caller — each should have agreed to this specific use.

A general appearance release is a start. It is often not enough on its own, because older releases rarely mention AI training or model building. Consent should reach the use you are actually selling.

You may not need a release for your own voice, or for people who are not identifiable, or for material where you already hold specific written consent covering training. The safe default is simple. If a person can be recognised, get their agreement in writing.

Excluding the people you cannot clear

You do not have to clear everyone. You can exclude. Cut the segment with the unconsented guest. Hold back the episode with the caller who never signed. License the rest.

This is why per-asset handling matters. When each clip is licensed on its own terms, a single unconsented voice does not taint a whole catalogue. You license what is clean and leave out what is not.

On fiund, consent is captured and checked at the asset level. Owners keep ownership, approve which buyers may license their material, and the platform records provenance so a buyer can see that the voices in a set were cleared. Where a person is identifiable, voice and likeness consent is handled separately from the content licence.

Watch outOld appearance releases often predate AI. A release that never mentions training or model building may not cover the use you are licensing now.
TipKeep a simple two-bucket system while recording: consent captured, and consent pending. Never license from the second bucket.

Sources

← All resources

Frequently asked questions

Do I need consent from someone who is only heard faintly in the background?

If they cannot be identified — an indistinct part of a crowd — the risk is low. If a listener could pick out and quote what one person said, treat them as identifiable and get consent or cut the clip.

I own the copyright in my podcast. Is that enough to license it for training?

It lets you license the recording, but not, by itself, the voices of identifiable guests. Their voice and likeness rights are separate and need their consent for AI-training use.

Can I just remove the parts I cannot clear?

Yes. Excluding an unconsented segment is often the cleanest path. Per-asset licensing means one uncleared voice does not stop you licensing the rest.

Related resources

Want data that clears this in diligence?

Whether you're building a model or sitting on an archive, the first conversation is short and specific.

Send a brief