Modality

Conversational speech for AI training

Real people talking to each other.

Conversational speech is unscripted, multi-party talk: two or more people negotiating meaning in real time, with interruptions, overlaps, backchannels, and mid-sentence repairs. It is the hardest speech to model and the hardest to buy. Much of what is sold as a conversational audio dataset is actually prompted dialogue — two crowd workers reading a scenario at each other. Real conversation behaves differently. Turn gaps are short, often a fraction of a second. Speakers talk over each other. Sentences get abandoned and restarted. Models trained only on read or acted speech degrade on exactly these features.

The scarcity has a legal root, not a technical one. Calls, meetings, and interviews are recorded constantly, but recording a conversation does not give anyone the right to license it. Every speaker on the channel has to have consented, and consent gathered after the fact is slow and expensive. That is why speech dataset licensing for genuine conversation is largely a sourced-to-brief business: the material has to be collected, or cleared, deliberately.

What separates a good corpus from a bad one: per-speaker channels or verified diarization, honest metadata (device, environment, relationship between speakers), transcripts that keep disfluencies instead of tidying them into prose, and consent paperwork that names AI training explicitly. A corpus missing any of these is cheaper for a reason.

Why it's scarce — and why that matters

Most speech corpora are scripted and read aloud. Genuine conversation — interruptions, crosstalk, accents, overlapping talk — is what speech and audio models are short on, and it is the supply fiund has the warmest path to.

Capture specs that matter

ASR pipelines typically consume 16 kHz mono, so capture above that — 44.1 or 48 kHz — preserves options for TTS and paralinguistic work later. Per-speaker close mics (headset or lavalier) beat a single far-field mic because they make overlap separable and give diarization ground truth; a far-field room channel is worth keeping alongside as a realistic test condition. SNR above roughly 20 dB is comfortable training material; below 10 dB, transcription cost and error rates climb fast. The annotations that matter: verbatim transcripts with disfluency marks, timestamped speaker turns (RTTM or JSON), overlap flags, and language tags wherever speakers code-switch.

Typical delivery formats: WAV, FLAC, JSON.

What it's good for

What drives licence cost

No two briefs price the same. These are the factors that move a Conversational speech licence up or down:

  • Speakers per recording — every additional voice is additional consent work
  • Spontaneity — genuine conversation costs more to source than acted dialogue
  • Transcription depth — verbatim with timestamps and speaker labels vs audio-only
  • Language, dialect, and accent rarity
  • Channel setup — matched close-talk plus far-field pairs command more than a single mixed track
  • Metadata depth — demographics, device, environment, speaker relationship
  • Exclusivity — a sole licence prices above non-exclusive

What to inspect before you licence

A sample and an hour of diligence catch most bad corpora. Check:

  • Pull random clips and listen for interruptions and backchannels — if nobody ever talks over anyone, it is acted dialogue
  • Verify per-speaker channels or diarization labels against the audio, not just the paperwork
  • Check transcripts keep disfluencies ("um", false starts) rather than cleaned-up prose
  • Confirm consent covers every voice on the recording, not just the account holder
  • Divide hours by unique speakers — 100 hours from 10 people is not 100 hours of variety
  • Ask for device and environment metadata, and check it actually varies

Rights & provenance

Every Conversational speech asset fiund lists carries a signed licence, explicit AI-training rights, and separate voice/likeness consent where people are identifiable. Nothing is scraped. Read more in the rights & provenance guides.

Related dataset specs

Frequently asked questions

How is conversational speech different from read speech for ASR training?

Read speech is fluent and predictable; conversation has disfluencies, overlap, fast turn-taking, and abandoned sentences. ASR systems trained mostly on read speech show markedly higher word error rates on spontaneous talk. The two are complements, not substitutes — most serious ASR briefs include both, weighted toward conversation.

Do I need per-speaker channels, or is a single mixed channel enough?

It depends on the task. Diarization and overlap-robust ASR benefit enormously from separated channels, because they provide ground truth for who spoke when. A single far-field channel is a realistic test condition but a weak sole training source. The strongest corpora ship both, time-aligned.

How should crosstalk and overlapping speech be handled?

It should be kept and labelled, not edited out. Overlap is a core feature of real conversation and a known failure mode for ASR and diarization systems. Good corpora mark overlapping segments and, where per-speaker channels exist, let you reconstruct each voice separately.

Can a conversation be licensed if only one participant consented?

No. Every identifiable voice on the recording needs to have consented to the licence, and where the speech makes a person identifiable, likeness and voice consent apply on top of copyright. fiund requires consent from all speakers before an asset lists — one-party call recordings do not clear that bar.

What transcription standard should I expect with a conversational corpus?

Verbatim transcription including disfluencies, timestamped speaker turns, and marked overlap and unintelligible regions. From that base you can force-align, simplify, or normalize downstream. A corpus delivered only with cleaned, readable transcripts has thrown away information you cannot recover.

Need Conversational speech data?

Send a brief and we source to spec, with the rights cleared before anything moves.

Send a brief