Read speechText-to-speech (TTS)

Read speech for Text-to-speech (TTS)

Text-to-speech (TTS) needs clean, consistent, studio-grade recordings with permissioned voices and matching transcripts. Here is why Read speech is the right raw material for it, what buyers typically spec, and how the rights are handled.

Why Read speech for Text-to-speech (TTS)

TTS is the use case read speech exists for. A synthesis model treats every sound in the corpus as the voice, so the properties that matter are consistency and consent, in that order. Consistency: fixed mic distance, one room, stable energy across sessions — drift between hour one and hour thirty becomes audible wobble in the synthesized voice. Coverage: phonemically balanced scripts so no phoneme in context is undertrained, plus enough prosodic range (questions, lists, emphasis) that the voice does not flatline. And consent: the output is a clone-adjacent artifact of a real person, so the licence must name synthesis, not just "training". Sample rate is the unforgiving spec — 44.1 or 48 kHz capture, because modern vocoders synthesize above what 16 kHz source material contains, and no post-processing recovers the missing band.

What buyers typically spec

Industry-typical ranges — a brief can and should deviate where the task demands it.

Typical volumeA few hours to tens of hours per voice depending on target quality; multi-voice corpora for base models
Audio44.1/48 kHz, 24-bit, treated room, fixed mic setup, consistent level
ScriptsPhonemically balanced prompts plus prosodic variety (questions, emphasis, numbers, names)
LabelsExact prompt text per utterance; misreads flagged; optional phoneme alignments
FormatsWAV/FLAC per utterance

A sample brief

The shape of a workable request — swap in your own numbers and conditions:

  • Modality: studio read speech, single professional voice, neutral plus expressive styles.
  • Volume: 15 hours over multiple sessions with matched setup.
  • Scripts: phonemically balanced set plus dialogue-style and long-form passages.
  • Rights: voice consent naming synthesis and cloning; exclusivity terms stated up front.

fiund's sourcing angle

Useful for TTS and pronunciation work, but only when the voices are permissioned. fiund sources read speech with explicit voice consent, which off-the-shelf corpora usually lack. We source to a brief and clear the rights before anything moves, so what you receive is both useful and defensible in diligence.

Rights posture

Signed licence, explicit training rights, separate voice/likeness consent, nothing scraped. See the rights & provenance guides.

← All Read speech data

Frequently asked questions

How consistent do sessions really need to be?

Audibly identical. Mic distance, room, and vocal energy drift are absorbed by the model as part of the voice, producing timbre wobble between phrases. The standard check is A/B listening between first and last sessions before accepting delivery.

Does the corpus need expressive styles or just neutral reads?

Depends on the product. Neutral-only corpora produce flat voices that cannot do emphasis or dialogue. If the voice will read anything beyond announcements, spec style coverage — questions, excitement, softness — as separate labelled subsets.

What consent language does a TTS voice need?

A grant that names training and voice synthesis explicitly, states exclusivity, and survives diligence. "Recorded for hire" or research-scoped releases are the classic gap: they cover the recording, not the cloned voice.

Other data for Text-to-speech (TTS)

More Read speech use cases

Need Read speech for Text-to-speech (TTS)?

Whether you're building a model or sitting on an archive, the first conversation is short and specific.

Send a brief