KarYeah · Egocentric data for embodied AI Sector 17, Chandigarh · 30.74°N 76.79°E
KarYeah
Chandigarh · India

Robots learn from the first person. We supply the eyes.

We record how people actually do things — hands, gaze, tools, mess and all — then parse and filter that footage into training-ready episodes for embodied AI.

Why first-person

Third-person video teaches a machine to watch. Egocentric video teaches it to act.

A robot arm never sees the scene from across the room. It sees what a hand sees: an object half-occluded by fingers, a handle at an odd angle, a surface that shifts as the head turns.

That viewpoint gap is why models trained on scraped web video stall the moment they touch real hardware. KarYeah closes it — recording from the body, keeping only what a policy can learn from.

POV · two-wheeler service bay, Mohali · wrist IMU on

What "parsed" means

One frame. Five signals pulled out of it.

hand · L
kettle
chai glass
Action segmentpour → chai glass · conf 0.94

The network, in motion.

380+ recorders · North India

Rec 01
Kitchen prepLudhiana
Rec 02
Two-wheeler repairMohali
Rec 03
Orchard pruningHimachal
Rec 04
Kirana counterChandigarh
Rec 05
Instrument prepPanchkula
Rec 06
Last-mile handoffZirakpur
Rec 01
Kitchen prepLudhiana
Rec 02
Two-wheeler repairMohali
Rec 03
Orchard pruningHimachal
Rec 04
Kirana counterChandigarh
Rec 05
Instrument prepPanchkula
Rec 06
Last-mile handoffZirakpur

The pipeline · in order

Four stages, from a kitchen in Mohali to your training bucket.

01

Collect

Trained recorders capture task episodes on calibrated head-mounted rigs. Every session starts with a signed consent form and a scene check.

02

Parse

Footage is synced, de-warped and cut into action segments with hand pose, gaze, object tracks and a plain-language narration per segment.

03

Filter

Blurred frames, dropped tracking, bystander faces and near-duplicate takes are removed. What survives is rare, clean and genuinely varied.

04

Deliver

Episodes are packed into LeRobot, RLDS or your schema, with a data sheet listing provenance, consent status and known gaps.

What survives the filter

We keep about a third. Drag to see the rest go.

Submitted 612 episodesShipped 188 episodes · keep rate 30.7%
0Hours captured
0Trained recorders
0Labelled action segments
0Kept after filtering

Figures updated quarterly · last revision Q2 2026

Ready to license

Collections built around real tasks, not benchmarks.

KY-012,900 hrs

Kitchen & canteen

Prep, pouring, plating and clean-up across home kitchens, dhabas and canteen lines. Dense two-hand manipulation with constant occlusion.

KY-023,400 hrs

Workshop & repair

Two-wheeler service, electrical repair and carpentry. Tool selection, fine alignment, and the failure-and-retry loops policies rarely see.

KY-032,100 hrs

Retail & warehouse

Picking, stacking, barcode handling and shelf resets in small shops and micro-fulfilment floors. Heavy on navigation plus grasp.

Browse all collections

Consent & provenance

Every frame can be traced to a person who agreed to record it.

Recorders are paid, briefed and free to pull a session back after the fact. Buyers get the paperwork, not just the pixels.

Read the data charter
Consent
Written, per session, in the recorder's own language. Withdrawable for 30 days after capture.
Bystanders
Faces, plates, screens and documents blurred automatically, then sample-checked by a reviewer.
Locations
Site permission on file for every commercial and workplace capture.
Provenance
Per-episode manifest: rig, recorder ID, date, city, consent reference, filter decisions.
Licensing
Non-exclusive or exclusive. No resale of client-commissioned captures.

Questions we get

Answers, before you email us.

What is egocentric data?

Egocentric data is recorded from a person's own point of view — usually head-mounted video, plus gaze, hand pose, audio and motion sensors. Because the camera moves with the body, it captures the same viewpoint a robot or wearable agent has to act from.

How do you protect people who appear in the footage?

Every recorder signs a consent agreement before capture. Faces, licence plates, screens and documents are blurred during the parse stage, and a human reviewer checks a sample of every batch before release. Recorders can withdraw a session for 30 days after it is filed.

What formats do you deliver?

Video as MP4 or MKV with per-frame timestamps, annotations as JSON or Parquet, and full episodes in LeRobot, RLDS or a schema you specify. Delivery by S3, GCS or encrypted physical drive.

Can you capture a task that isn't in your catalogue?

Yes. Commissioned collection is most of our work. Send the task list and the rig constraints; a scoped pilot of 40–100 episodes usually runs in three to four weeks.

Who is behind KarYeah?

KarYeah is a venture of 172 Tech, built and run from Chandigarh. The name is Punjabi shorthand — kar, to do. The site was designed by 530 Expert.

Tell us the task. We'll record it.

Request data Become a recorder