Egocentric videoAction recognition

Egocentric video for Action recognition

Action recognition needs labelled human activity across viewpoints and conditions. Here is why Egocentric video is the right raw material for it, what buyers typically spec, and how the rights are handled.

Why Egocentric video for Action recognition

Egocentric action recognition is its own discipline. Third-person models watch a body move; first-person models must infer action from hands, objects, and camera motion — the actor is mostly invisible. Class structure changes accordingly: fine-grained verb-object combinations ("cut tomato", "open drawer", "pour water") replace whole-body classes, and performance hinges on hand-object interaction cues. The camera moves with the head, so motion blur, sudden viewpoint shifts, and self-occlusion by the actor’s own arms are standard conditions, not defects — models trained on stabilized or third-person footage collapse here. Wearable products (AR glasses, assistive tech, industrial compliance) all consume exactly this distribution. Data value concentrates in annotation: tight temporal boundaries on short actions, consistent verb-noun taxonomies, and enough participants that hand appearance and technique vary.

What buyers typically spec

Industry-typical ranges — a brief can and should deviate where the task demands it.

Typical volumeTens of thousands of labelled action instances across fine-grained classes
VideoHead-mounted 1080p+ 30–60 fps; natural motion retained
LabelsVerb-noun class per segment, start/end boundaries, active-object tags
CoverageMany participants and environments; left/right handedness represented
FormatsMP4; JSON annotations

A sample brief

The shape of a workable request — swap in your own numbers and conditions:

  • Modality: first-person daily-activity footage (cooking, cleaning, assembly).
  • Volume: 100 verb-noun classes, 300+ instances each, 80+ participants.
  • Labels: temporal boundaries, active object, hand visibility flags.
  • Rights: participant consent naming AI training; shared-space bystander policy.

fiund's sourcing angle

Egocentric data is scarce and hard to collect at quality, and it is exactly what robotics and world-model teams need. fiund sources it to brief. We source to a brief and clear the rights before anything moves, so what you receive is both useful and defensible in diligence.

Rights posture

Signed licence, explicit training rights, separate voice/likeness consent, nothing scraped. See the rights & provenance guides.

← All Egocentric video data

Frequently asked questions

Why do third-person action models fail on egocentric footage?

The signal is different: no visible body, constant camera motion, actions evidenced by hands and object state changes. Viewpoint, scale, and motion statistics all shift, so egocentric performance requires egocentric training data — transfer from third-person is weak.

How fine-grained should the taxonomy be?

As fine as the product decision it feeds, and no finer — verb-noun classes multiply fast and thin out per-class data. Fix the taxonomy before collection so instances are gathered to fill it, rather than taxonomizing whatever arrived.

Other data for Action recognition

More Egocentric video use cases

Need Egocentric video for Action recognition?

Whether you're building a model or sitting on an archive, the first conversation is short and specific.

Send a brief