Motion captureGesture synthesis

Motion capture for Gesture synthesis

Gesture synthesis needs synchronized speech and motion capture from consented performers. Here is why Motion capture is the right raw material for it, what buyers typically spec, and how the rights are handled.

Why Motion capture for Gesture synthesis

Co-speech gesture models generate motion from audio, so the training asset is a synchronized pair: what was said, and what the body did while saying it. Both channels must be capture-grade — mocap for the motion, clean speech audio with transcripts for the conditioning — and the sync between them is the labour that matters, because gesture-speech timing operates at beat level; tens of milliseconds of drift blur the alignment the model exists to learn. Content requirements are specific: natural monologue and conversation rather than performed gesturing, enough hours per speaker to learn personal style (gesture is strongly idiosyncratic), and multiple speakers to span styles. Finger capture is near-mandatory — hands carry most communicative gesture, and body-only corpora produce arm-waving mittens. Emotional and discourse variety (explaining, disagreeing, storytelling) drives the gesture repertoire the model can produce.

What buyers typically spec

Industry-typical ranges — a brief can and should deviate where the task demands it.

Typical volumeHours per speaker for style modelling; multi-speaker breadth for general models
CaptureOptical mocap with finger capture, 60–120 Hz; time-locked close-mic audio
LabelsVerbatim transcripts with timestamps; discourse/affect tags per segment
SyncAudio-motion alignment verified and stated (sub-frame drift target)
FormatsBVH/FBX motion, WAV audio, JSON alignment

A sample brief

The shape of a workable request — swap in your own numbers and conditions:

  • Modality: synchronized mocap + speech of natural monologue and dialogue.
  • Volume: 3 hours each from 8 speakers, finger capture included.
  • Labels: timestamped transcripts, discourse-type tags per segment.
  • Rights: consent covering both voice and motion for AI training.

fiund's sourcing angle

Mocap studios own valuable libraries but lack a licensing channel with proper performer consent. fiund provides both. We source to a brief and clear the rights before anything moves, so what you receive is both useful and defensible in diligence.

Rights posture

Signed licence, explicit training rights, separate voice/likeness consent, nothing scraped. See the rights & provenance guides.

← All Motion capture data

Frequently asked questions

Can I combine separate speech and mocap corpora?

Not for this task — the model learns the coupling between specific words and specific movements, which only exists in simultaneously captured pairs. Unpaired data can pretrain each side, but the core corpus must be synchronized capture.

Why is finger capture near-mandatory here?

Because communicative gesture concentrates in the hands: points, counts, shapes, beats. Body-only capture reduces gesture to arm swing, and models trained on it generate exactly that. Budget for gloves or marker-based finger capture up front.

More Motion capture use cases

Need Motion capture for Gesture synthesis?

Whether you're building a model or sitting on an archive, the first conversation is short and specific.

Send a brief