Resources

Far-field and diarization audio data

Meeting and in-room speech is distant, overlapping, and multi-mic. Training robust ASR and diarization needs that acoustic reality, captured with consent.

Published 2026-07-22 · 6 min read

Key takeaways

  1. Far-field audio is defined by distance, reverberation, noise, and overlapping speech, not just content.
  2. Diarization (who spoke when) is scored by Diarization Error Rate; overlap and distance drive it up.
  3. Capture close-talking reference channels alongside distant mics to get clean ground truth.
  4. Every speaker is identifiable by voice, so per-speaker consent is a requirement, not an option.
  5. Custom capture lets you commission the exact rooms, arrays, and overlap your model is weak on.

Close-talking audio is the easy case. A headset near the mouth, one speaker, low noise. Real rooms are harder. Microphones sit meters away. People talk over each other. Reverberation, background noise, and moving speakers all degrade the signal. Systems that transcribe meetings and label who spoke have to work in that setting, which means they have to be trained and tested on it.

The research challenges make the difficulty concrete. The CHiME series targets distant, multi-microphone conversation; CHiME-6 used four-speaker dinner parties recorded on ear-worn mics and six microphone arrays. The AMI Meeting Corpus offers about 100 hours of meetings with close-talking and far-field microphones. These are the conditions custom capture has to reproduce.

What makes far-field hard

Distance is the first problem. As the microphone moves away from the speaker, the direct sound weakens and reflected sound grows. Reverberation smears words together. Background noise competes. The same sentence that transcribes cleanly on a headset can fall apart on a table microphone across the room.

Overlap is the second. In real conversation people interrupt and talk at once. Overlapping speech, sometimes called crosstalk, is where both transcription and speaker labeling break down. Benchmarks that once excluded overlap flattered their systems; modern ones, such as the DIHARD challenges, score it directly.

Diarization: who spoke when

Diarization answers a separate question from transcription: not what was said, but who said it and when. It segments audio by speaker and assigns each segment to a voice. Meeting transcription needs both, run together, so the transcript reads as a labeled conversation.

The field measures diarization with Diarization Error Rate. DER sums three failures: missed speech, false alarms, and speaker confusion, over the total reference speech. Overlap and distant microphones push all three up, which is why the acoustic conditions of the training data matter as much as its volume.

The acoustic conditions to specify

Far-field data is defined by its capture setup. Specify the number and placement of microphones, and whether arrays are used. State the room types and their reverberation, the number of simultaneous speakers, and the expected overlap. Note the noise environments, whether a quiet office or a busy cafe.

Reference signals decide what you can measure. Many programs capture a close-talking channel per speaker alongside the distant microphones. That gives a clean ground truth for transcription and speaker boundaries, which is hard to recover from the far-field mix alone.

Why custom consented capture beats scraped

Scraped meeting audio fails on two fronts. Technically, you inherit whatever microphones and rooms happened to be used, with no control over overlap, array geometry, or reference channels. You cannot commission the conditions your model is weak on.

On rights, recorded conversation is worse than most media. Every speaker is identifiable by voice, which is a likeness and, in several regimes, a biometric identifier. Recording people without consent, then training on their voices, is exactly the exposure buyers screen for. Scraped conversation carries no consent chain.

Sourcing far-field audio to a brief

A brief specifies the scenario: room, microphone layout, speaker count, languages, and overlap. It sets the reference channels and the annotation plan for speaker labels and transcripts. It fixes the acoustic difficulty you need rather than the difficulty you happened to find.

Sourced-to-brief capture records consenting speakers in the target conditions, with per-speaker consent for voice and likeness, and provenance attached. Contributors keep ownership and license the use for AI training. You get the hard acoustic cases, the reference signals to score them, and a rights position that survives diligence.

TipRecord a close-talking channel for each speaker at the same time as the distant microphones. It is the cleanest way to get ground-truth transcripts and speaker boundaries for scoring.
Watch outScraped meeting or call audio has no consent chain, and every voice in it is identifiable. Training on it is the kind of exposure buyer diligence is built to catch.

Far-field capture spec

  • Microphone count, placement, and array geometry
  • Room types and reverberation profile
  • Number of simultaneous speakers and expected overlap
  • Noise environments (office, cafe, vehicle)
  • Reference channels: close-talking mic per speaker
  • Per-speaker voice/likeness consent and provenance

Sources

← All resources

Frequently asked questions

What is the difference between far-field ASR and diarization?

Far-field ASR transcribes distant speech into words. Diarization labels who spoke and when. Meeting transcription runs both together, so the output is a conversation attributed to speakers.

Why not use close-talking data and add noise later?

Simulated reverberation and added noise help, but they do not fully reproduce real room acoustics, microphone arrays, or the way people overlap. Models trained only on augmented clean audio tend to degrade on genuine far-field recordings.

How many microphones should we specify?

It depends on the target device. A single distant microphone, a linear array, and a set of distributed arrays pose different problems. Specify the layout your product uses, and capture a reference channel per speaker for scoring.

Related resources

Want data that clears this in diligence?

Whether you're building a model or sitting on an archive, the first conversation is short and specific.

Send a brief