Roundup

Best speech data providers

Speech is the modality where rights questions bite hardest: a voice can be biometric data, and consent terms decide whether a corpus survives diligence. This roundup covers providers that actually deliver speech and voice data — crowd collectors recording to spec, catalog vendors with off-the-shelf hours, and boutique studios with human QA. The ranking weighs coverage and scale first, then how much rights documentation a buyer can actually inspect. fiund licenses owner-held recordings with per-speaker consent and is not ranked in its own roundup; where a vendor below is the better fit for your brief, its review says so.

  1. Defined.aiThe established stop for off-the-shelf speech and language corpora — a marketplace running since 2015, citing 1.6M+ contributors across 500+ languages and dialects. Confirm per asset that the licence grants AI-training use explicitly. fiund vs Defined.ai
  2. AppenThe scale option for commissioned collection: speech and audio gathered to spec in 500+ locales through the CrowdGen crowd. Rights live in each SOW, so consent terms are yours to set — and yours to check. fiund vs Appen
  3. ShaipHealthcare-speech depth few others carry: 70,000+ hours across 65+ languages, 250K hours of physician dictation, and 30M unstructured patient notes. Thinner outside its clinical and speech core. fiund vs Shaip
  4. NexdataThe volume play — company-cited holdings around 3M speech hours, with unsupervised pools of 100,000+ hours per language across seven major languages. Compliance is asserted at company level, so per-dataset consent records are the diligence item. fiund vs Nexdata
  5. PangeanicTwo decades of language-services work behind multilingual speech and ASR/TTS datasets, with regional programs for Arabic, Chinese, African, and Southeast Asian languages and sovereign-AI deployment options for European buyers. fiund vs Pangeanic
  6. Way With WordsThe boutique pick: commissioned speech with matched transcripts, multi-stage human QA, and consent documented back to the collection brief — plus rare practical depth in African-English and low-resource languages. fiund vs Way With Words
  7. The Vocal MarketNot speech but singing — consent-documented vocal stems (500+ recordings, 5,000+ stems in four languages) with an explicit AI-training opt-in. The option when your voice model needs vocals rather than conversation. fiund vs The Vocal Market

Other roundups

Let's talk about what you actually need.

Whether you're building a model or sitting on an archive, the first conversation is short and specific.

Send a brief