Motion capture → Gesture synthesis
Motion capture for Gesture synthesis
Gesture synthesis needs synchronized speech and motion capture from consented performers. Here is why Motion capture is the right raw material for it, what buyers typically spec, and how the rights are handled.
Why Motion capture for Gesture synthesis
Co-speech gesture models generate motion from audio, so the training asset is a synchronized pair: what was said, and what the body did while saying it. Both channels must be capture-grade — mocap for the motion, clean speech audio with transcripts for the conditioning — and the sync between them is the labour that matters, because gesture-speech timing operates at beat level; tens of milliseconds of drift blur the alignment the model exists to learn. Content requirements are specific: natural monologue and conversation rather than performed gesturing, enough hours per speaker to learn personal style (gesture is strongly idiosyncratic), and multiple speakers to span styles. Finger capture is near-mandatory — hands carry most communicative gesture, and body-only corpora produce arm-waving mittens. Emotional and discourse variety (explaining, disagreeing, storytelling) drives the gesture repertoire the model can produce.
What buyers typically spec
Industry-typical ranges — a brief can and should deviate where the task demands it.
| Typical volume | Hours per speaker for style modelling; multi-speaker breadth for general models |
|---|---|
| Capture | Optical mocap with finger capture, 60–120 Hz; time-locked close-mic audio |
| Labels | Verbatim transcripts with timestamps; discourse/affect tags per segment |
| Sync | Audio-motion alignment verified and stated (sub-frame drift target) |
| Formats | BVH/FBX motion, WAV audio, JSON alignment |
A sample brief
The shape of a workable request — swap in your own numbers and conditions:
- Modality: synchronized mocap + speech of natural monologue and dialogue.
- Volume: 3 hours each from 8 speakers, finger capture included.
- Labels: timestamped transcripts, discourse-type tags per segment.
- Rights: consent covering both voice and motion for AI training.
fiund's sourcing angle
Mocap studios own valuable libraries but lack a licensing channel with proper performer consent. fiund provides both. We source to a brief and clear the rights before anything moves, so what you receive is both useful and defensible in diligence.
Rights posture
Signed licence, explicit training rights, separate voice/likeness consent, nothing scraped. See the rights & provenance guides.
Frequently asked questions
Can I combine separate speech and mocap corpora?
Not for this task — the model learns the coupling between specific words and specific movements, which only exists in simultaneously captured pairs. Unpaired data can pretrain each side, but the core corpus must be synchronized capture.
Why is finger capture near-mandatory here?
Because communicative gesture concentrates in the hands: points, counts, shapes, beats. Body-only capture reduces gesture to arm swing, and models trained on it generate exactly that. Budget for gloves or marker-based finger capture up front.
More Motion capture use cases
Need Motion capture for Gesture synthesis?
Whether you're building a model or sitting on an archive, the first conversation is short and specific.
Send a brief