Collect
Calibrated head rigs, wrist IMUs and optional depth. Recorders work from a task script with a target count of episodes and a required variation list.
Raw first-person footage is close to useless — shaky, over-exposed, full of bystanders and repeated takes. The value is in what happens after the record button. Here is every stage, and what you get out of each.
Calibrated head rigs, wrist IMUs and optional depth. Recorders work from a task script with a target count of episodes and a required variation list.
Sync, de-warp, stabilise. Then hand pose, gaze projection, object tracks, contact events and a narration line per action segment.
Automatic rejection for motion blur, tracking loss, exposure blowout and near-duplicate takes; manual review for consent and framing.
Packed episodes plus a data sheet: provenance, consent references, filter statistics and the gaps we know about.
Capture stack
Filtering
A model does not improve because you fed it more hours. It improves because the hours were different from each other and clean enough to learn from.
Our filter stage scores every episode on tracking integrity, exposure, occlusion balance and novelty against everything already in the collection. Episodes that only repeat what the set already contains get dropped, no matter how well shot they are.
Delivery
LeRobot, RLDS, WebDataset shards, or a custom schema you define. Annotations as JSON or Parquet.
Your S3 or GCS bucket, an SFTP endpoint, or encrypted physical drives for large one-time transfers.
Per-collection documentation modelled on Datasheets for Datasets: how it was made, who by, what it under-represents.