Glossary
TTS (text-to-speech)
Technology that synthesizes natural-sounding speech from text; trained on clean, consistent, permissioned voice recordings.
A dedicated TTS voice is typically built from hours of clean, consistent recordings of one speaker; multi-speaker and zero-shot systems train on much larger, more varied corpora. Output quality is judged by human listening tests, reported as MOS. Because the output imitates a real voice, TTS training data needs voice consent on top of copyright clearance.
Why it matters
Clean, consented recordings of a single voice are among the most licensable audio assets an archive can hold.