Glossary

TTS (text-to-speech)

Technology that synthesizes natural-sounding speech from text; trained on clean, consistent, permissioned voice recordings.

A dedicated TTS voice is typically built from hours of clean, consistent recordings of one speaker; multi-speaker and zero-shot systems train on much larger, more varied corpora. Output quality is judged by human listening tests, reported as MOS. Because the output imitates a real voice, TTS training data needs voice consent on top of copyright clearance.

Why it matters

Clean, consented recordings of a single voice are among the most licensable audio assets an archive can hold.

See also

← All terms