Glossary
ASR (automatic speech recognition)
Technology that transcribes spoken language into text; trained on audio paired with accurate transcripts across accents and conditions.
ASR models train on paired audio and transcripts, from carefully aligned corpora to large weakly supervised collections. Accuracy is reported as word error rate and varies with accent, background noise, vocabulary, and speaking style — spontaneous conversation is consistently harder than read speech. That is why conversational, accented, and noisy recordings are the material ASR teams still seek.
Why it matters
ASR teams usually buy data to fix specific failure modes, so datasets described by accent, domain, and acoustic conditions match briefs faster.