Sung audio → Text-to-speech (TTS)
Sung audio for Text-to-speech (TTS)
Text-to-speech (TTS) needs clean, consistent, studio-grade recordings with permissioned voices and matching transcripts. Here is why Sung audio is the right raw material for it, what buyers typically spec, and how the rights are handled.
Why Sung audio for Text-to-speech (TTS)
Singing-voice synthesis is TTS with a second axis: the model must land the right phoneme on the right pitch at the right time. That makes the data requirement stricter than spoken TTS. The voice must come as dry stems — synthesis models trained on mixed masters learn reverb and compression as if they were vocal traits. The text must be time-aligned lyrics, ideally at phoneme level, because singing stretches syllables across notes in ways plain transcripts cannot express. And a pitch reference — f0 track or note-level MIDI — turns audio into supervised training pairs for pitch-conditioned synthesis. Range coverage matters the way phoneme coverage does in speech: a voice recorded only mid-range cannot be synthesized convincingly at its extremes. Consent is the sharpest issue in the category, since the output can imitate a performer’s identity.
What buyers typically spec
Industry-typical ranges — a brief can and should deviate where the task demands it.
| Typical volume | A few hours of one voice for voice-specific work; multi-singer corpora for base models |
|---|---|
| Audio | 44.1/48 kHz, 24-bit dry vocal stems; no reverb, tuning, or compression baked in |
| Labels | Time-aligned lyrics (word/phoneme), f0 or MIDI note track, section and tempo markers |
| Coverage | Full usable range, multiple dynamics and techniques per singer |
| Formats | WAV/FLAC stems; JSON/MIDI annotations |
A sample brief
The shape of a workable request — swap in your own numbers and conditions:
- Modality: solo vocal performances, dry stems, original or commissioned material.
- Volume: 5 hours per singer, 4 singers spanning ranges and styles.
- Labels: phoneme-aligned lyrics and f0 tracks; take-level technique notes.
- Rights: performer consent naming voice synthesis; composition rights documented.
fiund's sourcing angle
Music rights are famously tangled. fiund only lists sung audio where the performer and rights are documented — a small but genuinely clean supply. We source to a brief and clear the rights before anything moves, so what you receive is both useful and defensible in diligence.
Rights posture
Signed licence, explicit training rights, separate voice/likeness consent, nothing scraped. See the rights & provenance guides.
Frequently asked questions
Why do mixed masters fail for singing-voice synthesis?
The model cannot separate voice from production, so it learns reverb tails, compression pumping, and backing bleed as vocal characteristics. Dry stems are the training-grade asset; mixes are reference material at best.
What alignment granularity do lyrics need?
Phoneme-level is the working standard for synthesis, because singing decouples syllable duration from speech timing — one vowel may span many notes. Word-level alignment supports retrieval and rough conditioning but undershoots for generation.
Does consent for singing differ from spoken TTS consent?
Same structure, higher stakes: a singing voice is closer to a public identity, so the grant should state whether outputs may resemble the performer or only feed blended voices, and whether use is exclusive. Ambiguity here fails diligence.
Other data for Text-to-speech (TTS)
More Sung audio use cases
Need Sung audio for Text-to-speech (TTS)?
Whether you're building a model or sitting on an archive, the first conversation is short and specific.
Send a brief