Modality
Agentic trajectories for AI training
Recorded task demonstrations.
Agentic trajectories are recordings of humans actually completing tasks, captured at the level an agent can learn from: screen state plus every click and keystroke for digital work, or synchronized video, pose, and object state for physical work. Demand tracks the rise of computer-use agents and robot learning, and supply has not caught up — this is the youngest licensable modality on this list.
The scarcity is structural. The richest trajectory data is platform telemetry, and platforms do not sell it. Screen recordings are dense with personal and commercial information, so consent and redaction are heavy. And a trajectory is only useful if actions and observations are aligned tightly enough to reconstruct cause and effect — casual screen capture does not clear that bar. The result is that almost everything real is collected to a brief: a defined task list, purpose-built capture tooling, and consent designed in from the start.
Good corpora share traits: an explicit goal statement per episode, timestamped action logs aligned to observations, terminal-state labels (success, failure, abandoned), and — critically — retained failures and corrections. Demonstration sets scrubbed down to perfect runs teach an agent nothing about recovery, and recovery is most of the job.
Why it's scarce — and why that matters
Agentic training data barely exists as a licensable category yet. fiund treats it as a sourcing brief from day one.
Capture specs that matter
Digital trajectories pair screen video or DOM/accessibility-tree snapshots with an event log — clicks, keys, scrolls — timestamped on a single clock; JSON event streams plus MP4 or periodic screenshots are the common shape. Physical demonstrations add pose, object state, and sometimes force or contact signals. PII redaction has to happen without breaking action-observation alignment, which is why capture on controlled accounts beats after-the-fact scrubbing. Episode metadata: goal, app or environment version, outcome label, duration.
Typical delivery formats: JSON, CSV.
What it's good for
What drives licence cost
No two briefs price the same. These are the factors that move a Agentic trajectories licence up or down:
- Task expertise — specialist workflows (accounting, CAD, clinical) price above generic browsing
- Instrumentation depth — DOM plus video plus events vs screen video alone
- Redaction burden — live-account capture costs more to clean than sandboxed capture
- Episode diversity — distinct tasks and environments, not repeats
- Failure and recovery coverage — deliberately retained, not scrubbed
- Exclusivity — early-category buyers often want sole licences
What to inspect before you licence
A sample and an hour of diligence catch most bad corpora. Check:
- Replay sample episodes and verify logged actions reproduce the observed state changes
- Check timestamp alignment tolerance between events and frames
- Compare goal statements against what was actually done in the episode
- Confirm failure and correction episodes are present, not filtered out
- Audit PII redaction — complete, but with action-observation alignment intact
- Check environment versioning — an unversioned UI makes episodes unreproducible
Rights & provenance
Every Agentic trajectories asset fiund lists carries a signed licence, explicit AI-training rights, and separate voice/likeness consent where people are identifiable. Nothing is scraped. Read more in the rights & provenance guides.
Frequently asked questions
Is there a standard format for agent trajectories?
Not yet. The working shape is a JSON event log aligned to screen video or snapshots, but schemas differ by lab and tooling. Define the schema in the brief — action vocabulary, observation format, alignment tolerance — and expect light adaptation work on any corpus you did not commission.
Are failed episodes worth paying for?
Yes — arguably more than clean runs. Corrections and recoveries are where an agent learns what to do when a step misfires, and all-success corpora systematically overstate downstream performance. A corpus with zero failures was curated into unreality.
How is on-screen personal and commercial data handled?
The clean pattern is capture on controlled accounts and sandboxed environments so sensitive data never enters the recording; the fallback is a redaction pass that must not break event-to-frame alignment. Either way the demonstrator signs consent, and the protocol — not luck — is what keeps client data out.
Digital and physical trajectories — same brief?
No. Digital capture is software instrumentation; physical demonstration needs cameras, pose or mocap, and object tracking, closer to an egocentric-video collection. They share the episode/goal/outcome structure but not the capture stack, so scope them as separate briefs.
Need Agentic trajectories data?
Send a brief and we source to spec, with the rights cleared before anything moves.
Send a brief