Dataset spec · Sourced to brief

US English contact-centre conversations (sourced to brief)

This spec is sourced to brief: fiund sources the material directly from owners and clears the rights to your exact requirements before anything moves.

Task-oriented US English phone conversations with consent and PII handling in place, sourced to a buyer’s spec for voice-agent and ASR training.

At a glance

ModalityConversational speech
Use casesAutomatic speech recognition (ASR), Voice agents, Speaker diarization
Languagesen-US
Formatswav, json
Licenceai-training
Consent on fileYes
AvailabilitySourced to brief

Per-clip metadata

Every asset ships with a machine-readable manifest (JSON or CSV, mapped to your ingestion schema on request). Standard fields:

asset_idStable unique identifier, consistent across re-deliveries
durationHH:MM:SS.mmm
format / codecContainer and codec of the delivered file (originals preserved where licensed)
resolution · frame_rateVideo assets: source resolution and fps, no upscaling
sample_rate · channels · bit_depthAudio assets: as recorded; lossless masters where the owner has them
languageISO 639-1, per clip
speaker_count · diarizationSpeaker count and time-aligned speaker turns where processed
transcriptTime-aligned transcript where available or commissioned
category · tagsContent category and descriptive tags
recorded_at · country_of_originRecording date (where known) and ISO 3166 country
rights_refReference to the signed licence covering the asset — the chain-of-rights record producible in diligence
consent_statusWhether releases are on file for identifiable people in the recording

Provenance

Signed licence with explicit AI-training rights, voice and likeness consent where people are identifiable, nothing scraped. Provenance records available in diligence — see rights & provenance.

Request this dataset.

Send a brief and we scope it to your exact spec, rights cleared before anything moves.

Send a brief