Resources

Egocentric video datasets: a buyer’s guide

First-person wearable-camera video for robots, world models, and assistants. What to specify, and why rights are the hard part.

Published 2026-07-22 · 6 min read

Key takeaways

  1. Ego4D and Ego-Exo4D are benchmarks, not broad commercial training licences; check terms before you train.
  2. Egocentric quality is decided at capture: viewpoint, rig, sync, and annotations cannot be added later.
  3. The hard part is rights: bystanders, private spaces, faces, and voices all need a consent chain.
  4. For a shippable model, specify consented, sourced-to-brief capture with a signed AI-training licence.
  5. Keep provenance per clip so identifiable-person consent survives diligence.

Egocentric video is shot from the wearer’s point of view. A head- or chest-mounted camera sees what the person sees: their hands, the tools they hold, the people they talk to, the room around them. Labs training robots, world models, and wearable assistants want this footage because it matches how an embodied agent perceives the world.

Public research corpora set the reference points. Ego4D spans about 3,670 hours from 931 participants across 74 locations in nine countries. Ego-Exo4D adds time-synchronized first- and third-person views, roughly 1,422 hours from more than 800 participants in 13 cities. Both are strong for benchmarking. Neither is a general licence to train a commercial model on your own brief. This guide covers what to specify, and how consented capture fills the gap.

Why labs want first-person video

An embodied model has to act from a single, moving viewpoint. It sees hands enter and leave the frame. It watches objects get picked up, used, and put down. That is the signal egocentric video carries and third-person footage does not.

The Ego4D benchmark suite shows the tasks this data supports: episodic memory, hand-and-object interaction, audio-visual diarization, social understanding, and future-action forecasting. Ego-Exo4D adds paired exocentric views and expert commentary for skilled activities, from cooking to bouldering. For world models and assistants, the value is the tight link between what the wearer sees and what the wearer does next.

The public-dataset landscape

Ego4D and Ego-Exo4D are the anchors. Epic-Kitchens contributes about 100 hours of unscripted kitchen activity with dense action labels. Much recent capture runs on research glasses such as Project Aria, which record video alongside gaze, audio, and inertial signals.

Read the licences before you plan a training run. These corpora ship under research-oriented terms, with de-identification applied to faces and other identifiers. They are built to compare methods, not to grant broad commercial training rights over identifiable people captured in homes and workplaces. For benchmarking, use them. For a model you will ship, you usually need capture that was consented and licensed for that purpose.

What to specify in a brief

Egocentric quality is set at capture time. Decide these up front.

Tasks and activities: the specific behaviors you need, not just a scene. Viewpoint and rig: glasses versus chest mount, field of view, resolution, frame rate. Synchronization: which streams must be time-aligned, such as a second camera, multichannel audio, IMU, and eye gaze. Environments: kitchens, workshops, streets, offices. Annotations: narrations, hand and object boxes, action-segment boundaries, and any expert commentary. Coverage: languages, regions, and participant mix, so the set is not narrow.

Why consent and rights are hard here

First-person capture records private life. The camera enters homes and kitchens. It films bystanders who never agreed to appear. It catches faces, screens, documents, and licence plates. A clip scraped from the web carries none of the permissions a training run needs.

Identifiable people raise a second layer. A recognizable face or an audible voice is a likeness, and in several regimes a biometric identifier. Blur and masking reduce that exposure but also strip the hand, object, and interaction detail that made the footage worth capturing. You cannot annotate your way out of a missing consent chain.

How sourced-to-brief capture solves it

The alternative is to capture for the purpose. Consenting participants record the activities in the brief, in real settings, on a rig you specify. Bystander and private-space protocols are agreed before anyone presses record. Every clip is traceable to the person who made it.

The rights follow the footage. Contributors sign an AI-training licence. Where a person is identifiable, separate voice and likeness consent is captured. Provenance is documented so it survives diligence. Contributors keep ownership and license the use. That is the difference between footage you can benchmark on and footage you can train and ship on.

Watch outResearch corpora such as Ego4D apply de-identification and ship under research-oriented licences. Do not assume they grant commercial rights to train on identifiable people captured in private settings.
TipAsk for IMU and eye-gaze streams time-aligned to video from the start. Retro-fitting synchronization to footage that was not captured with it is rarely possible.

Egocentric capture spec

  • Activities and tasks named, not just scenes
  • Rig and viewpoint: glasses or chest, field of view, resolution, frame rate
  • Synchronized streams listed: second camera, audio, IMU, gaze
  • Environments and any private-space protocol
  • Annotation plan: narrations, hand/object boxes, action segments
  • Consent per participant; voice/likeness consent where identifiable
  • Provenance recorded per clip

Sources

← All resources

Frequently asked questions

Can we train a commercial model on Ego4D or Ego-Exo4D?

Treat them as benchmarks first. Both ship under research-oriented licences with de-identification, so read the terms for your intended use. For a model you plan to ship, consented capture licensed for AI training is the cleaner path.

Glasses or chest-mounted camera?

It depends on the task. Glasses track head and gaze direction and sit close to the wearer’s line of sight. Chest mounts are steadier and less intrusive for long sessions. Specify the one that matches your target device.

How do you handle bystanders in public spaces?

With protocols agreed before capture: consented framing, private-space rules, and de-identification where a bystander cannot consent. Provenance records what was agreed for each clip.

Related resources

Want data that clears this in diligence?

Whether you're building a model or sitting on an archive, the first conversation is short and specific.

Send a brief