Modality

Motion capture for AI training

Consented performance, rig-ready.

Motion capture is movement recorded as data: joint rotations and positions over time rather than pixels. It drives avatar animation, humanoid robotics, pose estimation, and gesture models. The supply problem is unusual. Thousands of hours of high-grade mocap already exist in studio libraries, captured for games and film — and almost none of it can be licensed onward, because performer agreements were scoped to a production, not to AI training. The studios know it. So the libraries sit there, valuable and unsellable, while teams re-capture material that already exists somewhere.

Licensing solves it from the performer side: consent that names model training, payment attached to the licence, and a documented chain from performer to file. That is a paperwork product as much as a data product, and it is why clean mocap supply stays thin.

Buyers should care about skeleton discipline above almost everything else. A corpus captured across sessions with inconsistent skeletons, marker sets, or coordinate conventions costs real engineering weeks to unify. Good corpora ship one skeleton with a documented joint hierarchy, calibration per session, raw plus cleaned takes, and honest notes on what was filtered — over-aggressive cleanup strips the micro-motion that makes human movement read as human.

Why it's scarce — and why that matters

Mocap studios own valuable libraries but lack a licensing channel with proper performer consent. fiund provides both.

Capture specs that matter

Optical marker systems capture at 100–240 Hz with millimetre-order precision and remain the ground-truth standard; IMU suits trade accuracy for portability and drift; markerless multi-camera capture is improving but struggles with occlusion and contact moments. Common interchange formats: BVH and FBX for rig-ready skeletal animation, C3D for raw marker trajectories; SMPL-family parameters are increasingly requested in research pipelines. Confirm skeleton definition, frame rate, and whether finger and face capture are included — body-only is the default.

Typical delivery formats: BVH, CSV, JSON.

What it's good for

What drives licence cost

No two briefs price the same. These are the factors that move a Motion capture licence up or down:

  • Capture system grade — optical marker vs IMU suit vs markerless
  • Performer consent and skill — stunt, dance, and sign-language specialists price higher
  • Finger and face capture — separate setups, real added cost
  • Cleanup level — raw, cleaned, or animation-polished takes
  • Multi-person and prop interactions — harder to capture, rarer to find
  • Format and retargeting deliverables — BVH/FBX/C3D, plus custom skeletons
  • Exclusivity of takes or of a performer’s library

What to inspect before you licence

A sample and an hour of diligence catch most bad corpora. Check:

  • Load samples into your own rig and confirm the joint hierarchy matches the documentation
  • Check for foot sliding and ground penetration in cleaned takes
  • Look for marker-swap artifacts — sudden joint flips mid-take
  • Verify frame rate is consistent across the corpus, not mixed capture sessions
  • Ask for calibration records and T-pose/range-of-motion takes per session
  • Read the performer consent — production-scoped releases do not cover training

Rights & provenance

Every Motion capture asset fiund lists carries a signed licence, explicit AI-training rights, and separate voice/likeness consent where people are identifiable. Nothing is scraped. Read more in the rights & provenance guides.

Related dataset specs

Frequently asked questions

BVH, FBX, or C3D — what should I ask for?

BVH or FBX if you want rig-ready skeletal animation; C3D if you want raw marker trajectories to solve against your own skeleton; SMPL-family output if your pipeline is research-oriented. The safest corpora keep raw marker data so any future format can be derived. Ask what the studio archived, not just what they export.

Marker-based or markerless — does it matter for training data?

Marker-based optical capture is still the accuracy ceiling and the right ground truth for pose estimation. Markerless capture is cheaper and scales, but degrades on occlusion, contact, and fast motion. The honest framing: markerless is fine when mocap is your training volume, not when it is your ground truth.

Does motion capture include hands and face by default?

No. Body capture is the default; fingers need added markers or gloves, and faces need a separate facial capture setup. Gesture and avatar briefs that need hand articulation should say so explicitly — retrofitting finger data onto body-only takes is not possible.

Why can’t existing game and film mocap libraries just be licensed?

Because the performers signed for a production, not for AI training, and voice/likeness-adjacent rights in a performance do not transfer by default. Relicensing means going back to performers for consent — which is exactly the work a licensing marketplace does up front.

What frame rate does training data actually need?

Many animation and gesture targets train at 30–60 Hz, but dynamics work — robotics, physics models, sports — benefits from 100 Hz and above. Capture high and downsample; you cannot recover inter-frame motion later.

Need Motion capture data?

Send a brief and we source to spec, with the rights cleared before anything moves.

Send a brief