Modality

Sensor & IMU for AI training

Wearable and device sensor streams.

Sensor and IMU data is movement measured directly: accelerometer, gyroscope, and magnetometer streams from phones, wearables, and instrumented devices, usually paired with labels for what the wearer was doing. It trains activity recognition and health features, and increasingly feeds robotics and world-model work as a proprioceptive signal alongside video.

The paradox is abundance without access. Billions of devices produce these streams constantly, but the data pools inside platforms under privacy terms that do not permit resale — and stripping identity out of motion data is harder than it sounds, because gait is close to a fingerprint. So licensable supply means deliberate, consented collection: a participant, a device, a protocol, and labels.

Labels are where corpora live or die. Raw IMU streams without ground truth are nearly worthless; the value is the annotation — which activity, when it started, when it ended, and how that was verified. Good corpora document the labelling protocol (self-report, observer, or video-verified), keep all timestamps on one clock, state sampling rates per channel, and are honest about device placement: a model trained on wrist-worn data does not transfer cleanly to a phone in a pocket. Device, placement, and population diversity are what a brief should spec first.

Why it's scarce — and why that matters

Sensor data is abundant on devices but almost never licensed cleanly. fiund sources consented streams to spec.

Capture specs that matter

Consumer IMUs typically sample at 50–200 Hz: 3-axis accelerometer and gyroscope (6-axis), plus magnetometer for 9-axis. Higher rates matter for impacts and fine motion; 50 Hz is adequate for coarse activity classes. Timestamps drift across devices, so the corpus should state its sync method. Deliverables are usually CSV or JSON with per-channel rates, device model, placement (wrist, pocket, chest, ankle), and labelled activity spans. Video-verified labels are the quality ceiling; unverified self-report is the floor.

Typical delivery formats: CSV, JSON.

What it's good for

What drives licence cost

No two briefs price the same. These are the factors that move a Sensor & IMU licence up or down:

  • Label verification method — video-verified ground truth prices above self-report
  • Device and placement diversity across the corpus
  • Participant diversity — age, body type, mobility range
  • Protocol complexity — free-living collection costs more than scripted sessions
  • Multi-sensor synchronization — IMU aligned with video or depth
  • Activity rarity — falls, industrial tasks, and clinical movement are hard to collect

What to inspect before you licence

A sample and an hour of diligence catch most bad corpora. Check:

  • Plot random windows and sanity-check units (g vs m/s²) and the gravity axis
  • Check sampling-rate consistency and scan for gaps and dropped samples
  • Verify label boundaries against reference video where it exists
  • Divide hours by participants — many hours from few bodies overfits
  • Check class balance across activities, not just total hours
  • Confirm device model and placement metadata per session

Rights & provenance

Every Sensor & IMU asset fiund lists carries a signed licence, explicit AI-training rights, and separate voice/likeness consent where people are identifiable. Nothing is scraped. Read more in the rights & provenance guides.

Frequently asked questions

Scripted sessions or free-living collection — which do I need?

Scripted protocols give clean, dense labels but unnaturally crisp transitions; free-living data has realistic transitions and ambiguity but sparser, weaker labels. Most production briefs want a scripted core plus a free-living validation slice. Specify the split rather than buying one and hoping.

Can motion data really be anonymized?

Not reliably. Gait and movement patterns are strongly identifying, so treating stripped-ID sensor data as anonymous is a compliance bet. The defensible route is participant consent that names AI training — which is why consented collection, not anonymization, is the basis for licensing here.

What sampling rate should the brief specify?

50 Hz covers coarse activities like walking, sitting, and cycling. Impacts, sports, falls, and fine manipulation want 100–200 Hz or more. Capture at the higher rate when in doubt — downsampling is free, upsampling is fiction.

Does IMU data need to be paired with video?

For training, not necessarily — but video is the strongest label-verification tool, and multimodal briefs increasingly want synchronized IMU-plus-video so the sensor stream can supervise or be supervised by vision. Sync quality then becomes part of the spec.

Need Sensor & IMU data?

Send a brief and we source to spec, with the rights cleared before anything moves.

Send a brief