Resources

Sourcing data for world models and physical AI

World models learn to predict what happens next. They need long, continuous real-world video tied to actions and sensors, with clean provenance.

Published 2026-07-22 · 6 min read

The landscape here moves quickly — re-verify against the cited sources for the current status before you rely on it.

Key takeaways

  1. World models need continuity, actions, and sensor state, not just clips.
  2. The field pattern: broad video to learn dynamics, a smaller action-linked set to make the model controllable.
  3. Time alignment between video, actions, and sensors is a hard requirement; drift corrupts the signal.
  4. Real environments carry people and property; per-clip provenance and consent are the fix.
  5. Specify continuity, sensors, and diversity in the brief, because they cannot be retro-fitted.

A world model predicts what happens next. Given what an agent sees and does, it forecasts the next observation. That prediction is what lets a robot or an assistant plan instead of react. Training one takes more than clips. It takes long, continuous recordings of the real world, paired with the actions and sensor signals that explain the changes on screen.

The scale is visible in recent systems. NVIDIA Cosmos was trained on a curated pool of about 20 million hours of video, filtered toward motion, manipulation, and navigation. Meta’s V-JEPA 2 pre-trained on over one million hours of video, then learned action-conditioned prediction from under 62 hours of real robot interaction. The pattern is consistent: broad video to learn how the world moves, then a smaller, precise set of action-linked data to make the model controllable.

What a world model needs from data

Three things. Continuity: unbroken sequences long enough to show cause and effect, not two-second cuts. Actions: what the agent or person did, aligned in time to the video, so the model links a move to its result. State: the sensor signals that ground the pixels, such as inertial data, depth, camera pose, and force.

Internet video teaches general dynamics. It rarely comes with actions or calibrated sensors. That gap is why action-conditioned training uses a separate, smaller corpus of interaction data, and why teams increasingly commission it rather than hope to find it.

Long, continuous video

Length matters because physics plays out over time. A pour, a handover, a turn through a doorway: each needs a continuous take to show the setup, the motion, and the outcome. Hard cuts and montage break the signal.

Specify duration per take, frame rate, resolution, and camera motion. Egocentric and fixed-camera footage answer different questions; many world-model programs want both. State the scene diversity you need so the model does not overfit to one room or one lighting condition.

Actions and sensors

Action labels turn video into a controllable model. For robots, that is the joint or end-effector trajectory. For people, it can be hand pose, tool use, or step boundaries. Aggregated robot corpora such as Open X-Embodiment show the appetite: over a million trajectories pooled across many robot types.

Sensor streams ground the model in physical state. Inertial measurement gives motion and orientation. Depth and camera pose give geometry. Time alignment is the requirement that is easy to underestimate: a sensor log that drifts from the video by a fraction of a second corrupts the very link you are trying to learn.

Provenance and consent for real environments

World-model data is recorded in real places: homes, warehouses, streets, vehicles. Those settings contain people, private property, and, on public roads, other parties who never opted in. A large video pool with no per-clip provenance is a liability an AI lab has to inherit.

Consent and provenance are the fix. Capture in controlled or consented environments. Record who agreed to what, per clip. Where people are identifiable, treat their likeness and voice as separate consents. Keep the chain documented so it survives the buyer’s diligence.

Sourcing to a brief

A brief converts a model gap into a capture plan. Name the interactions, the environments, the sensors, and the alignment tolerances. Decide continuity and diversity before capture, because neither can be added afterwards.

Sourced-to-brief capture returns footage that was consented and licensed for AI training, with sensors synchronized and provenance attached. Contributors keep ownership and license the use. For physical AI, where the model will act in the world, that rights posture is not a formality. It is what lets you deploy what you trained.

NoteAction-conditioned training often uses far less data than pre-training. V-JEPA 2 learned control from under 62 hours of robot interaction after pre-training on more than a million hours of video.
Watch outA large video pool with no per-clip provenance transfers risk to the buyer. Diligence teams increasingly reject data that cannot show where each clip came from.

World-model capture spec

  • Continuous take length and frame rate
  • Action labels: trajectory, hand pose, or step boundaries
  • Sensor streams: IMU, depth, camera pose, force
  • Time-alignment tolerance across streams
  • Scene and lighting diversity
  • Consent and per-clip provenance for real environments

Sources

← All resources

Frequently asked questions

How is a world model different from a video generator?

A video generator produces plausible footage. A world model predicts the next state given an action, so an agent can plan against it. The difference in data is action conditioning: paired records of what was done and what happened next.

Do we need robots to collect action data?

Not always. Robot trajectories are one source. Human demonstrations with hand pose, tool use, and step labels are another, and egocentric capture supports them. The requirement is that actions are recorded and time-aligned to the video.

Why not just scrape long videos from the web?

Scraped video rarely has actions, calibrated sensors, or a consent chain. It can teach general dynamics but not control, and it carries provenance risk. Commissioned capture fills the action-linked gap and clears the rights.

Related resources

Want data that clears this in diligence?

Whether you're building a model or sitting on an archive, the first conversation is short and specific.

Send a brief