Roundup
Best speech data providers
Speech is the modality where rights questions bite hardest: a voice can be biometric data, and consent terms decide whether a corpus survives diligence. This roundup covers providers that actually deliver speech and voice data — crowd collectors recording to spec, catalog vendors with off-the-shelf hours, and boutique studios with human QA. The ranking weighs coverage and scale first, then how much rights documentation a buyer can actually inspect. fiund licenses owner-held recordings with per-speaker consent and is not ranked in its own roundup; where a vendor below is the better fit for your brief, its review says so.
- Defined.ai — The established stop for off-the-shelf speech and language corpora — a marketplace running since 2015, citing 1.6M+ contributors across 500+ languages and dialects. Confirm per asset that the licence grants AI-training use explicitly. fiund vs Defined.ai →
- Appen — The scale option for commissioned collection: speech and audio gathered to spec in 500+ locales through the CrowdGen crowd. Rights live in each SOW, so consent terms are yours to set — and yours to check. fiund vs Appen →
- Shaip — Healthcare-speech depth few others carry: 70,000+ hours across 65+ languages, 250K hours of physician dictation, and 30M unstructured patient notes. Thinner outside its clinical and speech core. fiund vs Shaip →
- Nexdata — The volume play — company-cited holdings around 3M speech hours, with unsupervised pools of 100,000+ hours per language across seven major languages. Compliance is asserted at company level, so per-dataset consent records are the diligence item. fiund vs Nexdata →
- Pangeanic — Two decades of language-services work behind multilingual speech and ASR/TTS datasets, with regional programs for Arabic, Chinese, African, and Southeast Asian languages and sovereign-AI deployment options for European buyers. fiund vs Pangeanic →
- Way With Words — The boutique pick: commissioned speech with matched transcripts, multi-stage human QA, and consent documented back to the collection brief — plus rare practical depth in African-English and low-resource languages. fiund vs Way With Words →
- The Vocal Market — Not speech but singing — consent-documented vocal stems (500+ recordings, 5,000+ stems in four languages) with an explicit AI-training opt-in. The option when your voice model needs vocals rather than conversation. fiund vs The Vocal Market →
Other roundups
Let's talk about what you actually need.
Whether you're building a model or sitting on an archive, the first conversation is short and specific.
Send a brief