Modality
Read speech for AI training
Clean, scripted, studio-grade voice.
Read speech is the controlled end of the speech spectrum: a known script, a quiet room, a consistent voice. Nothing about the audio is hard to produce. What is hard is the paperwork behind it. Read-speech corpora mostly get used to build voices — TTS, voice cloning, pronunciation models — and that makes the speaker’s consent the real product. A voice model built on a corpus without explicit synthesis rights is a liability with good audio quality.
This is where speech dataset licensing gets specific. “Recorded for hire” is not the same as “licensed for AI training”, and neither automatically covers cloning a voice. Older corpora were collected before voice synthesis was the obvious end use, so their releases rarely mention it. Buyers now ask for consent language that names training and synthesis, and for the right to audit it in diligence. Supply that meets that bar is thin.
Quality separates on boring details: script design (phonemically balanced prompts versus random sentences), recording consistency across sessions, noise floor, and transcript fidelity. A good corpus sounds identical on hour one and hour forty. A bad one drifts — different mic distance, different room, different energy — and the drift shows up directly in the synthesized voice.
Why it's scarce — and why that matters
Useful for TTS and pronunciation work, but only when the voices are permissioned. fiund sources read speech with explicit voice consent, which off-the-shelf corpora usually lack.
Capture specs that matter
TTS work wants 44.1 or 48 kHz capture at 24-bit in a treated room; 16 kHz is enough for ASR but throws away what synthesis needs. Standard setup is a large-diaphragm condenser at fixed distance with a pop filter and a low, stable noise floor. Scripts should be phonemically balanced when lexicon coverage is the goal. Deliverables: per-utterance WAV or FLAC, the exact prompt text (with misreads flagged), session metadata, and pronunciation notes for names and loanwords.
Typical delivery formats: WAV, FLAC.
What it's good for
What drives licence cost
No two briefs price the same. These are the factors that move a Read speech licence up or down:
- Voice exclusivity — a voice licensed to one buyer prices far above a shared voice
- Consent depth — synthesis and cloning rights price above analysis-only use
- Recording grade — treated studio vs home setup
- Hours per single voice — long consistent single-speaker sessions are operationally hard
- Language, dialect, and accent rarity
- Script design — custom prompt lists cost more than generic passages
What to inspect before you licence
A sample and an hour of diligence catch most bad corpora. Check:
- Compare the first and last session of a voice for drift in level, distance, and tone
- Check the noise floor in silences — hum, HVAC, and room tone end up in the synthesized voice
- Verify audio matches the prompt text exactly, with misreads flagged rather than silently kept
- Ask for a phoneme coverage report if pronunciation or lexicon work is the goal
- Read the consent language — it should name synthesis or cloning, not just "research"
- Listen for clipping and over-processing; denoiser artifacts get learned by TTS models
Rights & provenance
Every Read speech asset fiund lists carries a signed licence, explicit AI-training rights, and separate voice/likeness consent where people are identifiable. Nothing is scraped. Read more in the rights & provenance guides.
Frequently asked questions
What sample rate should I insist on for TTS data?
44.1 or 48 kHz at 24-bit. Modern neural TTS is commonly trained at 22.05–48 kHz output, and you cannot upsample your way out of a 16 kHz source. If the corpus will only ever feed ASR, 16 kHz is acceptable — but capturing high and downsampling keeps both doors open.
Does a work-for-hire recording contract cover AI training?
Not reliably. Work-for-hire settles who owns the recording, but voice cloning implicates the speaker’s voice and likeness beyond copyright, and older contracts predate synthesis as a use. Diligence teams now look for consent that names training and synthesis explicitly. That is the standard fiund papers to.
How many hours of one voice does a TTS build need?
It depends on the target. Classic single-speaker builds used tens of hours; fine-tuning a modern multi-speaker model can need far less, and zero-shot systems trade per-voice hours for corpus breadth. State the model and target quality in the brief and spec hours from there, rather than defaulting to a round number.
Why does session consistency matter so much?
Because a TTS model treats everything it hears as the voice. If mic distance, room, or energy shifts between sessions, the model averages the variants or oscillates between them. Consistency is the cheapest quality lever in read speech, and the first thing to audit.
Need Read speech data?
Send a brief and we source to spec, with the rights cleared before anything moves.
Send a brief