Egocentric video → World models
Egocentric video for World models
World models needs long, continuous, real-world video with motion and depth cues. Here is why Egocentric video is the right raw material for it, what buyers typically spec, and how the rights are handled.
Why Egocentric video for World models
World models learn to predict what happens next, and egocentric video is the richest licensable stream of embodied cause and effect: a viewpoint that moves through space, hands that change object state, environments that persist and get revisited. The properties that matter are temporal ones. Long, unbroken takes let a model learn that rooms stay where they were; clip compilations destroy that structure. Real head motion provides egomotion signal — and paired IMU makes it explicit, giving models a proprioceptive channel alongside pixels. Post-processing is the silent killer: stabilization, aggressive exposure smoothing, and cuts all falsify the dynamics being modelled. The ideal corpus reads like unedited experience — hours-long sessions, natural task flow, documented camera intrinsics, synchronized sensors — closer to a log of embodiment than a video product.
What buyers typically spec
Industry-typical ranges — a brief can and should deviate where the task demands it.
| Typical volume | Hundreds to thousands of hours; session length is a first-class spec (30+ min takes) |
|---|---|
| Video | Head-mounted, unstabilized, fixed exposure where possible; intrinsics per device |
| Extras | Synchronized IMU strongly preferred; gaze/depth where hardware allows |
| Labels | Light: activity/location metadata, revisit annotations; prediction needs footage more than labels |
| Formats | MP4; sensor streams as CSV/JSON |
A sample brief
The shape of a workable request — swap in your own numbers and conditions:
- Modality: continuous first-person recordings of daily routines, 30–90 min takes.
- Volume: 500 hours, 60+ participants, homes and workplaces revisited across sessions.
- Extras: time-aligned IMU; camera intrinsics per unit.
- Rights: commercial AI-training consent; household-member consent handled.
fiund's sourcing angle
Egocentric data is scarce and hard to collect at quality, and it is exactly what robotics and world-model teams need. fiund sources it to brief. We source to a brief and clear the rights before anything moves, so what you receive is both useful and defensible in diligence.
Rights posture
Signed licence, explicit training rights, separate voice/likeness consent, nothing scraped. See the rights & provenance guides.
Frequently asked questions
Why are long takes non-negotiable for world models?
Prediction over minutes requires training sequences over minutes: object permanence, environment persistence, and task structure only exist in continuous footage. Once chopped to short clips, that supervision is unrecoverable.
How much does synchronized IMU add?
A lot for the price: it disambiguates egomotion from scene motion, provides a clean action-conditioning signal, and anchors dynamics learning. If hardware supports it, sync at collection time — retrofitting alignment later is misery.
Other data for World models
More Egocentric video use cases
Need Egocentric video for World models?
Whether you're building a model or sitting on an archive, the first conversation is short and specific.
Send a brief