Resources

How to license your data to OpenAI, Anthropic, and other labs

How lab intake works, why mid-market owners rarely get in directly, and the path that does work.

Published 2026-07-22 · 8 min read

The landscape here moves quickly — re-verify against the cited sources for the current status before you rely on it.

Key takeaways

  1. Some labs publish intake forms, but they open for scale and clean rights.
  2. Mid-market catalogs often sit below the size for a direct lab deal.
  3. Aggregation through a marketplace is the realistic path in.
  4. Labs check licence, consent, provenance, and spec every time.
  5. A documented, brief-matched offer beats a large, undocumented one.

You have data a frontier lab would want. That does not mean you can email one and sell it. Intake at the largest labs is narrow, and it is built for scale, not for one catalog at a time.

Here is how lab intake actually works, why mid-market owners rarely get in through the front door, and the path that does work.

How lab intake works

Some labs publish an intake channel. OpenAI, for example, announced a Data Partnerships program that invites organisations to submit interest in contributing data across text, images, audio, and video. It describes two paths, an open-source archive and private datasets, and says it is looking for large-scale data that expresses human intention and is not already easy to find online. It also says it is not seeking sensitive personal information or third-party data it has no right to use.

Not every lab runs an equivalent open form. Other labs, including Anthropic, have moved toward licensed sourcing as legal pressure on undocumented training data has grown, and large data deals are generally negotiated directly or through intermediaries. Treat any public form as a front door that opens for a narrow set of very large or very specific offerings.

Why mid-market owners rarely get in directly

Labs optimise for breadth, volume, and clean rights. A direct data partnership carries overhead: legal review, consent verification, provenance checks, and integration. That overhead is worth it to a lab when the supply is large or rare. For a mid-sized studio or archive, the catalog is often too small on its own to justify a direct relationship, however good it is.

This is not a judgement on quality. It is a threshold problem. The data can be exactly what a lab needs and still sit below the size where a direct deal makes sense for them.

The marketplace and broker path

Aggregation is how mid-market supply reaches labs. A marketplace pools many owners, standardises the AI-training licence and the consent layer, verifies provenance, and presents curated, rights-cleared supply that a buyer can evaluate in one place. The owner clears the threshold by being part of a larger, consistent pool.

This is the path fiund is built for. Owners keep ownership and approve buyers, licences are non-exclusive by default, and provenance is available in diligence, so a lab sees standardised rights rather than a stack of one-off contracts.

What labs check before they buy

Whichever path you take, the checks are the same. Is there a signed AI-training licence on every asset. Is there consent for identifiable people. Can provenance be traced. Does the data meet a spec: modality, volume, language, quality. Prepare those answers before you approach anyone, because you will be asked for all of them.

How to prepare your catalog

Inventory what you hold. Confirm you have, or can obtain, voice and likeness consent for identifiable people. Attach metadata. Then respond to briefs rather than pitching everything at once. A tight, documented offer that matches a stated need travels much further than a large, undocumented one.

The licensed market is growing, and buyers are under more pressure than before to prove their training data was lawfully sourced. Documentation is now the difference between a catalog a lab can buy and one it cannot.

NoteUse official channels only. Where a lab publishes an intake form, that is the route. Be wary of anyone claiming private contacts or guaranteed placement at a named lab.

Sources

← All resources

Frequently asked questions

Can I just contact a lab directly?

You can try where a lab publishes an intake form, and it is worth doing if your catalog is large or rare. For most mid-sized owners, a direct deal is unlikely because the volume does not clear the lab’s threshold. Aggregation is the more realistic route.

Do labs want data with people in it?

Only with consent. Public intake guidance from labs tends to steer away from sensitive personal information and third-party data. If people are identifiable in your data, you need voice and likeness consent for training use before it is licensable.

Is non-exclusive licensing a problem for labs?

Usually not. Non-exclusive is a common default, and many buyers accept it. Exclusivity is a separate, priced decision. Do not give it away to get a foot in the door.

Related resources

Want data that clears this in diligence?

Whether you're building a model or sitting on an archive, the first conversation is short and specific.

Send a brief