Conversational speechSpeaker diarization

Conversational speech for Speaker diarization

Speaker diarization needs multi-speaker recordings with overlap, crosstalk, and turn-level labels. Here is why Conversational speech is the right raw material for it, what buyers typically spec, and how the rights are handled.

Why Conversational speech for Speaker diarization

Diarization — who spoke when — is trained and scored on exactly the material scripted corpora cannot provide: multiple voices sharing a channel, interrupting, and overlapping. Diarization error rate on clean, turn-taking speech looks solved; on real meetings and calls, overlapping speech remains the dominant error source, and systems only learn it from data where overlap genuinely occurs. Conversational speech recorded with per-speaker close mics plus a mixed room channel is the gold arrangement: the separated channels yield exact turn boundaries and overlap ground truth, while the mixed channel is the realistic input a deployed system faces. Speaker counts matter too — two-party calls, four-way meetings, and eight-person dinners are different problems. A diarization corpus should state its speaker-count distribution and overlap percentage, not just hours.

What buyers typically spec

Industry-typical ranges — a brief can and should deviate where the task demands it.

Typical volumeTens to hundreds of hours; diarization needs breadth of sessions more than raw hours
AudioPer-speaker close-talk channels plus a far-field mixed channel, time-aligned
LabelsTurn-level speaker segments (RTTM), overlap regions, speaker IDs consistent across a session
Key statisticsSpeaker count per session (2–8+), overlap percentage, turn-duration distribution
FormatsWAV/FLAC multi-channel; RTTM or JSON segment files

A sample brief

The shape of a workable request — swap in your own numbers and conditions:

  • Modality: natural multi-party conversations, 3–6 speakers per session.
  • Volume: 100 hours across 200+ distinct sessions and speaker groups.
  • Capture: lapel mic per speaker plus one far-field array channel, sample-aligned.
  • Labels: RTTM turns with overlap marked; session-level speaker metadata.

fiund's sourcing angle

Most speech corpora are scripted and read aloud. Genuine conversation — interruptions, crosstalk, accents, overlapping talk — is what speech and audio models are short on, and it is the supply fiund has the warmest path to. We source to a brief and clear the rights before anything moves, so what you receive is both useful and defensible in diligence.

Rights posture

Signed licence, explicit training rights, separate voice/likeness consent, nothing scraped. See the rights & provenance guides.

← All Conversational speech data

Frequently asked questions

Do I really need per-speaker channels for diarization data?

For ground truth, effectively yes. Human-annotated turns on a single mixed channel drift at boundaries and miss short backchannels; separated channels give you sample-accurate reference segmentation, from which the mixed channel becomes a perfectly labelled training input.

What overlap percentage should the corpus have?

Match your deployment. Casual multi-party conversation commonly runs meaningful double-digit overlap; formal meetings less; call-centre audio less still. A corpus should report its measured overlap rate so you can compose the mix your product will face.

How many distinct speakers does diarization training need?

More sessions and more distinct voices beat more hours per voice. Diarization learns to separate unseen speakers, so a corpus of few groups recorded at length generalizes worse than many shorter sessions with fresh speaker combinations.

More Conversational speech use cases

Need Conversational speech for Speaker diarization?

Whether you're building a model or sitting on an archive, the first conversation is short and specific.

Send a brief