Sung audio → Emotion & prosody
Sung audio for Emotion & prosody
Emotion & prosody needs labelled emotional speech across speakers and contexts. Here is why Sung audio is the right raw material for it, what buyers typically spec, and how the rights are handled.
Why Sung audio for Emotion & prosody
Singing is emotion made deliberate: vibrato depth, dynamics, breathiness, timing rubato — the acoustic levers of affect, performed on purpose and to an extreme spoken corpora rarely reach. That makes sung audio a dense training signal for models that read or generate expressive audio. It stretches the feature space: a model that has only heard conversational arousal has never seen the sustained high-energy phonation or controlled pianissimo that singing supplies. Labelling differs from speech emotion work: useful units are phrases and sections rather than turns, and annotation can lean on musical structure — dynamics markings, key, tempo — as weak labels alongside human ratings of expressive intent. The standing caveat: performed emotion is stylized. Corpora should label expressive technique and intended affect separately, so models do not learn that vibrato equals sadness.
What buyers typically spec
Industry-typical ranges — a brief can and should deviate where the task demands it.
| Typical volume | Tens of hours, phrase-labelled; expressive range prioritized over scale |
|---|---|
| Audio | 44.1/48 kHz dry stems; multiple takes of the same material at different expressive intents |
| Labels | Phrase-level expressive intent, valence/arousal ratings, technique tags (vibrato, belt, breathy) |
| Coverage | Singers, genres, and languages varied; restrained and extreme expression both present |
| Formats | WAV/FLAC; JSON labels |
A sample brief
The shape of a workable request — swap in your own numbers and conditions:
- Modality: solo vocal phrases performed at controlled expressive intents.
- Volume: 20 hours; same material sung at contrasting dynamics and affect.
- Labels: intent tags, technique tags, 3-rater valence/arousal per phrase.
- Rights: performer consent naming AI training; original material only.
fiund's sourcing angle
Music rights are famously tangled. fiund only lists sung audio where the performer and rights are documented — a small but genuinely clean supply. We source to a brief and clear the rights before anything moves, so what you receive is both useful and defensible in diligence.
Rights posture
Signed licence, explicit training rights, separate voice/likeness consent, nothing scraped. See the rights & provenance guides.
Frequently asked questions
Is performed emotion a valid training signal or a bias?
Both, depending on labels. Performance stylizes affect, which is fine when the corpus labels expressive technique and intended emotion separately — the model learns the mapping. Unlabelled, it teaches shortcuts like vibrato-means-sad.
How does labelling sung emotion differ from speech?
The unit is the musical phrase, not the conversational turn, and structure (dynamics, tempo, key) provides usable side information. Rater agreement is still imperfect, so per-rater labels and agreement statistics remain best practice.
Other data for Emotion & prosody
More Sung audio use cases
Need Sung audio for Emotion & prosody?
Whether you're building a model or sitting on an archive, the first conversation is short and specific.
Send a brief