Methodology

How we source and clear data.

Provenance isn't a checkbox at the end. It's the first thing we do.

In AI training, data provenance is the documented answer to three questions: where did this material come from, who owns it, and who agreed to what. Buyers' counsel ask for it because they have to. Owners should care because it's what makes an archive worth licensing at all.

1. We go to the owner

We source non-public material directly from the creators, studios, and archives who hold it — not from crawls or scrapes, and never from behind a login. Going to the owner is what makes everything downstream possible: you can't paper consent with a crawler, and you can't fix a sourcing shortcut later with a disclaimer.

2. We paper the rights before anything moves

Clearing means three things, each in writing. A signed licence granting AI-training rights explicitly — not implied from a terms-of-service. Voice and likeness consent handled separately wherever a person is identifiable. And chain of title checked, so the party granting rights actually holds them. Consent, likeness, and training rights are handled up front, not retroactively — retroactive clearance is where deals go to die.

3. The provenance record travels with the file

Every asset carries a provenance record: who the source is, what they own, the signed licence and its scope, the consent records mapped to the identifiable people, when and how the material was collected, and a file-level manifest. That's data provenance as a working document rather than a marketing word — it's what a buyer's diligence team actually reads. The full anatomy is in provenance records, explained.

4. Why chain of title decides deals

The deals that fall apart don't usually fall apart over quality — quality problems can be fixed. They fall apart because somewhere in the chain, someone couldn't show they had the right to grant what they were granting. Chain of title is the sequence of agreements connecting the person in the recording to the party signing the licence. If it's intact, a deal can close; if it's broken, no volume of hours or resolution saves it. That's why we check it before buyer match, not during closing — see chain of title for training data.

Frequently asked questions

What is a provenance record?

The documentation that travels with every asset: source and ownership, the signed licence and its scope, consent records mapped to the identifiable people, collection dates and method, and a manifest of the files. It is the artefact that turns “trust us” into “read it yourself.”

Do you ever use scraped or public-crawl data?

No. Everything is sourced from the people who own it, under a signed licence, never from crawls and never from behind a login. Scraped material cannot carry the consent and chain of title that make data provenance mean anything.

What happens when consent can’t be verified?

The material sits out. Episodes, sessions, or footage without verifiable consent are excluded from the licensable set — they are not waved through, and they don’t quietly ride along with the clean material.

I’m a buyer — can I see the provenance before I license?

Yes; that is what it exists for. Provenance documentation is available in diligence, and the full record travels with delivery, so your counsel reads the same paperwork we papered.

Related reading

Let's talk about what you actually need.

Whether you're building a model or sitting on an archive, the first conversation is short and specific.

Send a brief