Sung audioEmotion & prosody

Sung audio for Emotion & prosody

Emotion & prosody needs labelled emotional speech across speakers and contexts. Here is why Sung audio is the right raw material for it, what buyers typically spec, and how the rights are handled.

Why Sung audio for Emotion & prosody

Singing is emotion made deliberate: vibrato depth, dynamics, breathiness, timing rubato — the acoustic levers of affect, performed on purpose and to an extreme spoken corpora rarely reach. That makes sung audio a dense training signal for models that read or generate expressive audio. It stretches the feature space: a model that has only heard conversational arousal has never seen the sustained high-energy phonation or controlled pianissimo that singing supplies. Labelling differs from speech emotion work: useful units are phrases and sections rather than turns, and annotation can lean on musical structure — dynamics markings, key, tempo — as weak labels alongside human ratings of expressive intent. The standing caveat: performed emotion is stylized. Corpora should label expressive technique and intended affect separately, so models do not learn that vibrato equals sadness.

What buyers typically spec

Industry-typical ranges — a brief can and should deviate where the task demands it.

Typical volumeTens of hours, phrase-labelled; expressive range prioritized over scale
Audio44.1/48 kHz dry stems; multiple takes of the same material at different expressive intents
LabelsPhrase-level expressive intent, valence/arousal ratings, technique tags (vibrato, belt, breathy)
CoverageSingers, genres, and languages varied; restrained and extreme expression both present
FormatsWAV/FLAC; JSON labels

A sample brief

The shape of a workable request — swap in your own numbers and conditions:

  • Modality: solo vocal phrases performed at controlled expressive intents.
  • Volume: 20 hours; same material sung at contrasting dynamics and affect.
  • Labels: intent tags, technique tags, 3-rater valence/arousal per phrase.
  • Rights: performer consent naming AI training; original material only.

fiund's sourcing angle

Music rights are famously tangled. fiund only lists sung audio where the performer and rights are documented — a small but genuinely clean supply. We source to a brief and clear the rights before anything moves, so what you receive is both useful and defensible in diligence.

Rights posture

Signed licence, explicit training rights, separate voice/likeness consent, nothing scraped. See the rights & provenance guides.

← All Sung audio data

Frequently asked questions

Is performed emotion a valid training signal or a bias?

Both, depending on labels. Performance stylizes affect, which is fine when the corpus labels expressive technique and intended emotion separately — the model learns the mapping. Unlabelled, it teaches shortcuts like vibrato-means-sad.

How does labelling sung emotion differ from speech?

The unit is the musical phrase, not the conversational turn, and structure (dynamics, tempo, key) provides usable side information. Rater agreement is still imperfect, so per-rater labels and agreement statistics remain best practice.

Other data for Emotion & prosody

More Sung audio use cases

Need Sung audio for Emotion & prosody?

Whether you're building a model or sitting on an archive, the first conversation is short and specific.

Send a brief