Modality
Sung audio for AI training
Vocal performance and music.
Sung audio carries more rights in a single file than any other modality here. A recorded song can stack a master right (the recording), a composition right (the song), and the performer’s voice — three claims, often three different owners, sometimes three different contracts. Licensing it for AI training means clearing every layer, and one missing signature poisons the whole asset. That is why genuinely clean sung corpora barely exist at scale, and why the clean supply is mostly original or commissioned performance rather than commercial catalog.
For model builders, the interesting asset is usually the dry vocal stem: the voice alone, before reverb, compression, and mix. A mixed master teaches a model the production as much as the voice. Stems with time-aligned lyrics, pitch, and phrase timing are the difference between a singing-voice dataset and a pile of songs.
Good corpora also document the performer, not just the file: vocal range, style, language, and whether consent covers synthesizing a voice that sounds like theirs. Voice cloning concerns are sharpest in music — artists have names and fans — so buyers should expect the consent language to address synthesis directly, and should walk away from corpora where it does not.
Why it's scarce — and why that matters
Music rights are famously tangled. fiund only lists sung audio where the performer and rights are documented — a small but genuinely clean supply.
Capture specs that matter
44.1 kHz/24-bit is the floor; 48 or 96 kHz capture is common in studio work. Dry, unprocessed vocal stems are worth far more than mixed masters — effects and mastering are artifacts a generative model will learn. Watch for headphone bleed of backing tracks into the vocal mic; it can smuggle uncleared material into a "clean" stem. Useful annotations: time-aligned lyrics, phoneme timing, f0 or MIDI pitch tracks, section markers (verse, chorus), and tempo. A cappella takes recorded to a click keep alignment options open.
Typical delivery formats: WAV, FLAC.
What it's good for
What drives licence cost
No two briefs price the same. These are the factors that move a Sung audio licence up or down:
- Rights layers cleared — master, composition, and performer consent together vs recording-only
- Stems vs mixed masters — isolated dry vocals price above mixes
- Annotation depth — aligned lyrics, phonemes, and pitch tracks add real cost
- Performer rarity — range, technique, genre, language
- Synthesis rights — cloning-grade consent prices above analysis-only
- Exclusivity of the performance or the voice
What to inspect before you licence
A sample and an hour of diligence catch most bad corpora. Check:
- Get the rights map in writing: who holds the master, the composition, and the performer consent
- Solo the stems and confirm they are dry — reverb tails and compression pumping mean processing baked in
- Listen for click or backing-track bleed in headphone spill
- Spot-check lyric and pitch alignment on random phrases
- Confirm performer consent names voice synthesis if that is the use
- Check range, style, and language metadata against the audio, not the catalog sheet
Rights & provenance
Every Sung audio asset fiund lists carries a signed licence, explicit AI-training rights, and separate voice/likeness consent where people are identifiable. Nothing is scraped. Read more in the rights & provenance guides.
Frequently asked questions
Why can’t I license songs from a stock or production music library?
Production-music licences cover sync, broadcast, and reproduction — not model training — and many libraries now exclude machine learning explicitly. Beyond the licence text, the composition and master are usually held by different parties, and the performer’s consent to voice modelling was never asked. Training use needs its own grant across all layers.
Should I ask for stems or full mixes?
Stems, if voice is what you are modelling. Mixed masters entangle the voice with instruments, reverb, and mastering, and models trained on them reproduce the production along with the performance. Mixes are only preferable when full-song structure is itself the training target.
What annotations matter for singing-voice synthesis?
Time-aligned lyrics at word or phoneme level, a pitch track (f0 or MIDI), and phrase or section boundaries. Singing couples text to melody, so text-only transcripts undershoot — the model needs to know which note carried which syllable.
How does performer consent differ from spoken-voice consent?
The mechanics are the same — a named person granting training and synthesis rights — but the stakes are higher because a singing voice is closer to a public identity. Consent should state whether outputs may imitate the performer’s voice or only contribute to blended voices. Ambiguity here is a diligence flag.
Need Sung audio data?
Send a brief and we source to spec, with the rights cleared before anything moves.
Send a brief