Dataset 03 · Human world model
Egocentric manipulation demonstrations
10,000 hrs · 20,000 clipsPer-hand labels · 27 industriesSupervised · SFT
First-person, head-mounted video of real workers on real job sites — 10,000 hours of it, labeled frame by frame with what each hand is doing, the tool it holds and the scene around it. Nothing is scraped from the internet: every clip describes what a person can actually do through their behavior, not a written claim. Demonstration data for imitation learning — the supervision a vision-language-action model or manipulation policy learns from directly.
- Footage
- 10,000 hours · 20,000 clips
- People
- 15,003 workers · 1,463 sites
- Coverage
- 27 industries, front-line long tail
- Volume
- 1.99 TB, frame-aligned
- Video
- Egocentric MP4 · H.264 + fps / exposure
- Annotation
- Left / right hand — object, action, intent
- Sensors
- 6-axis IMU (accel + gyro), frame-synced
- Rights
- Signed consent · faces & PII redacted
per-clip package:
snippets/{uuid}.mp4 egocentric video · H.264 + fps / exposure
hands[].{side,span,text} left / right action tracks, frame-aligned
imu/{uuid}.parquet 6-axis accel + gyro, synced to video
LICENSE · consent.csv signed consent · redaction · rights chain