Modality
Egocentric video for AI training
First-person, wearable-camera footage.
Egocentric video is filmed from the doer’s point of view — a head-mounted or chest-mounted camera capturing hands, tools, and workspace as a task unfolds. It is the modality robotics and world-model teams ask for most, because the viewpoint matches what an embodied agent actually sees: hands entering the frame, objects being grasped and released, attention shifting toward the next step.
It is scarce for a plain reason: it does not occur naturally. People do not wear head cameras through their day, so there is no archive to clear and nothing meaningful to scrape. Every hour has to be deliberately collected — hardware on a consenting person, a task worth recording, and enough sessions to cover variation in environment and technique. Academic corpora exist, but their licences typically stop at research use, which puts commercial teams back at square one.
Good corpora are defined by what surrounds the pixels. Camera intrinsics and mounting position. Synchronized IMU or gaze where the hardware supports it. Task and step annotations. Long unbroken takes rather than trimmed highlights. And restraint in post-processing: motion blur and camera shake are part of the signal, and a corpus that stabilized them away has destroyed information an agent learning from human video needs.
Why it's scarce — and why that matters
Egocentric data is scarce and hard to collect at quality, and it is exactly what robotics and world-model teams need. fiund sources it to brief.
Capture specs that matter
Typical capture is 1080p–4K at 30–60 fps on head-mounted action cameras or camera glasses, with wide-FOV lenses whose distortion should be documented (intrinsics per device). Synchronized IMU — and gaze, on devices that support it — multiplies value for embodied learning. Useful annotations: task and step segmentation, hand and object boxes or masks, spoken or written step narration, and camera mounting metadata. Long continuous takes preserve the temporal context that clip collections lose.
Typical delivery formats: MP4.
What it's good for
What drives licence cost
No two briefs price the same. These are the factors that move a Egocentric video licence up or down:
- Collection is bespoke — participant time and hardware are real per-hour costs
- Task complexity and environment access — industrial sites price above kitchens
- Annotation depth — step labels, hand-object masks, narration
- Sensor sync — IMU, gaze, or depth alongside video
- Participant and environment variety
- Exclusivity — commissioned collections are often sole-licensed
What to inspect before you licence
A sample and an hour of diligence catch most bad corpora. Check:
- Confirm lens FOV and distortion parameters are documented per device
- Check footage was not stabilized or aggressively post-processed
- Verify hands are actually visible during manipulation, not cropped by the mount angle
- Look at the session-length distribution — trimmed highlight reels lose temporal context
- Test video-IMU sync accuracy on sample takes if sensor streams are included
- Check consent covers bystanders appearing in shared spaces
Rights & provenance
Every Egocentric video asset fiund lists carries a signed licence, explicit AI-training rights, and separate voice/likeness consent where people are identifiable. Nothing is scraped. Read more in the rights & provenance guides.
Related dataset specs
Frequently asked questions
Why can’t I just use academic egocentric datasets commercially?
Most ship under research-oriented or restricted licences, and the participant consent behind them was scoped to research. Terms vary, so read them — but commercial training and deployment usually needs separately licensed collection with consent that names commercial use.
Head-mounted or chest-mounted — which should the brief specify?
Head mounts capture a gaze-aligned view with more motion; chest mounts are steadier and keep hands centred but miss where attention goes. Robotics briefs often prefer head-mounted for the attention signal; some manipulation work prefers chest. Specify it — the two are not interchangeable in training.
Do I need IMU or gaze sync, or is video alone enough?
Depends on the target. Visual pretraining and action recognition run fine on video alone. World models and embodied-learning work benefit materially from synchronized IMU, and gaze is a strong attention signal where hardware provides it. Sync adds collection cost, so gate it on the use.
Is camera shake a defect in egocentric footage?
No — it is signal. Head motion encodes locomotion, attention shifts, and task rhythm. Corpora that stabilize or heavily filter footage to look better have removed information embodied models train on. The defect to reject is missing calibration, not visible motion.
Need Egocentric video data?
Send a brief and we source to spec, with the rights cleared before anything moves.
Send a brief