A producer reviewing conversation recordings and audio tracks.

FROM THE SOURCE

A clearer
perspective.

Practical guides to the decisions behind AI data: the recordings, the rights, and the details that make a collection useful.

Human at the source.

FROM THE SOURCE

Ideas you can put to work.

Featured

For buyers · 8 min read

How to license real-world media for AI training

Start with the model task, then license the content, people, and delivery rights that task actually requires. This is the practical path from brief to documented dataset.

For buyers · 8 min read

How to buy training data with commercial AI rights

Commercial AI rights come from a written, traceable licence - not a download button. Learn which permissions, recipients, restrictions, and records to verify before production use.

For buyers · 8 min read

UGC data licensing for AI models

User-generated content can support vision, video, speech, and multimodal models, but public availability does not clear the creator, participants, music, brands, locations, or AI use.

For buyers · 8 min read

Alternatives to web scraping for training data

Licensed archives, commissioned capture, opt-in contributor programs, first-party data, partnerships, and hybrid synthetic programs can replace or narrow scraping while improving provenance.

For sellers · 5 min read

How to sell your data to AI without losing ownership

Licensing is not a buyout. How to keep your copyright, authorize fiund to handle AI licensing, withdraw from future deals, and get paid per licence.

Buyers & sellers · 9 min read

How AI data licensing works, end to end

The full path from brief to audit: sourcing, consent, manifest, licence, delivery, and payment.

For buyers · 6 min read

Egocentric video datasets: a buyer’s guide

First-person wearable-camera video for robots, world models, and assistants. What to specify, and why rights are the hard part.

Buyers & sellers · 7 min read

What’s in an AI data licensing agreement? A clause-by-clause guide

The clauses that decide a data deal — the grant, scope, derivative-model rights, warranties, indemnity, and takedown — in plain terms.

For sellers · 6 min read

Non-rival data: license the same recordings to many models

A recording is not used up when a model trains on it. That single fact is why non-exclusive licensing multiplies what an owner can earn.

For buyers · 7 min read

Licensed vs scraped training data: what’s safe to train on

The web is public, so scraping feels free. The law is more complicated. Here is what the current cases actually decided, and what safe to train on means for a buyer.

Buyers & sellers · 6 min read

What content can be licensed for AI training?

A practical map of audio, video, motion capture, LiDAR, sensor streams, and task demonstrations — plus the rights and documentation each category needs.

For sellers · 6 min read

How to make money creating content for AI training

How creators, studios, archives, and collection operators can prepare original material for AI programs — without promising income or giving up ownership by default.

Sell & license your data

For owners of audio, video, and archives who want to license to AI teams and keep their rights.

For sellers · 5 min read

How to sell your data to AI without losing ownership

Licensing is not a buyout. How to keep your copyright, authorize fiund to handle AI licensing, withdraw from future deals, and get paid per licence.

For sellers · 4 min read

What makes your recordings valuable to AI buyers

Buyers pay for what a crawl cannot give them. The qualitative drivers that make a recording worth licensing.

For sellers · 5 min read

How to license a video archive to AI teams

Clear a film or stock archive asset by asset: rights status, releases for people on camera, provenance, and deliverables.

For sellers · 5 min read

Licensing call-centre and customer-call recordings to AI

Turn customer calls into licensable speech data without skipping consent, redaction, and the law over recorded voices.

For sellers · 5 min read

Licensing research interviews and field recordings to AI

License interviews and field recordings the honest way: participant consent, ethics alignment, de-identification, ownership.

For sellers · 6 min read

How to license your voice to AI safely

License your voice without signing it away. Bounded terms, revocation, and the laws now written for voice cloning.

For sellers · 5 min read

Licensing spoken-word archives (lectures, sermons, courses) to AI

Prepare decades of lectures, sermons and courses for licensing: who owns the recording, and speaker consent at scale.

For sellers · 6 min read

Non-rival data: license the same recordings to many models

A recording is not used up when a model trains on it. That single fact is why non-exclusive licensing multiplies what an owner can earn.

For sellers · 5 min read

How AI training royalties work vs a one-time buyout

When you license data for AI training, you usually face one choice first: an ongoing licence or a one-time buyout. Here is how each model works and what you trade.

For sellers · 6 min read

How to make money creating content for AI training

How creators, studios, archives, and collection operators can prepare original material for AI programs — without promising income or giving up ownership by default.

Buying training data

For AI teams deciding what to license, build, or commission.

For buyers · 8 min read

How to license real-world media for AI training

Start with the model task, then license the content, people, and delivery rights that task actually requires. This is the practical path from brief to documented dataset.

For buyers · 8 min read

How to buy training data with commercial AI rights

Commercial AI rights come from a written, traceable licence - not a download button. Learn which permissions, recipients, restrictions, and records to verify before production use.

For buyers · 8 min read

UGC data licensing for AI models

User-generated content can support vision, video, speech, and multimodal models, but public availability does not clear the creator, participants, music, brands, locations, or AI use.

For buyers · 8 min read

Alternatives to web scraping for training data

Licensed archives, commissioned capture, opt-in contributor programs, first-party data, partnerships, and hybrid synthetic programs can replace or narrow scraping while improving provenance.

For buyers · 6 min read

How much speech data do you need to train an ASR model?

The honest answer is a range tied to your goal. Here are the tiers, the drivers, and why hours only matter alongside a sourcing and rights plan.

For buyers · 6 min read

How much audio do you need to train or fine-tune a TTS voice?

Cloning, fine-tuning, and building a voice need very different amounts of audio. And because a voice is a likeness, consent is part of the spec.

For buyers · 7 min read

Licensed vs scraped training data: what’s safe to train on

The web is public, so scraping feels free. The law is more complicated. Here is what the current cases actually decided, and what safe to train on means for a buyer.

Buyers & sellers · 5 min read

The data not in the crawl

Models have already read the open web. The signal that moves them now is the material that never entered a crawl. Here is why not-in-the-crawl data is worth sourcing.

For buyers · 6 min read

Buy vs build vs commission training data

Buy, build, or commission your training data. Most comparisons stop at cost and speed. The lens that decides it is legal defensibility.

For buyers · 6 min read

Synthetic vs real training data: when real wins

Synthetic data is cheap and endless. It is also a copy of what a model already knows. Here is where synthetic works, and where licensed real data is the only fix.

Frontier modalities

Sourcing egocentric video, agentic trajectories, world-model, and far-field data.

Content types for AI training

Audio, video, motion capture, LiDAR, sensors, and task demonstrations — what each category needs before it can be licensed.

Diligence & quality

How to vet a dataset, a vendor, and the rights behind them.

Rights & consent

Consent, releases, and who can license what.

For sellers · 6 min read

Can I sell data that has other people in it?

You own the recording. Other people are in it. Here is when that is fine, when you need a release, and when to leave a clip out.

For sellers · 6 min read

Podcast guest release forms for AI training

A standard appearance release rarely covers model training. Here is what to add, and why forward consent beats chasing it later.

Buyers & sellers · 6 min read

What consent for AI training actually looks like

Consent that stands up is specific, informed, and auditable — and it separates the content licence from the voice behind it.

For sellers · 6 min read

BIPA and your voice recordings: what owners should know

Your voice can count as biometric data. Illinois BIPA is why. If you record or license audio of identifiable people, consent at the source protects you as much as the buyer.

For sellers · 6 min read

Who owns the rights to your podcast?

You made the podcast, so you own it, right? Usually not all of it. Here is who holds which rights, with a worked example, before you license an episode for AI training.

Buyers & sellers · 6 min read

The US state voice and likeness law map

There is no single US law for AI voice and likeness. There is a growing patchwork of state statutes and a pending federal bill. Here is the map as it stands.

For buyers · 7 min read

EU AI Act data requirements: a buyer’s checklist

If you place a general-purpose AI model on the EU market, you have to publish a summary of what you trained on. Here is what Article 53 requires, and a checklist to get your data ready.

Deal terms

How AI data licences are structured.

How it works

The mechanics of an AI data licensing deal, end to end.