Use Case
Robotics & Manipulation
Robots learn manipulation best from human demonstrations. Payana turns real-world, first-person footage of people doing tasks into robot-ready training episodes - the imitation-learning data behind robotics foundation models.
Egocentric (head- or body-mounted) video is the closest human analogue to what a robot's own camera sees: the workspace, the objects, and the hands manipulating them, all from the actor's point of view. Payana captures this footage with wearable cameras while workers simply do their jobs - no scripting, no staging, no lab setup.
Each recording flows through an automated episode pipeline: a vision-language model narrates the action, detection and open-vocabulary models find every object and tool, segmentation produces pixel-accurate masks, speech becomes transcripts, and hand-pose estimation extracts a 7-dimensional end-effector state per frame - position, orientation, and gripper aperture, the same action space robot policies are trained in.
The result is demonstration data that pairs observation with action at every timestep - what the person saw and what their hands did - which is precisely the supervision imitation learning and behavior-cloning policies need. Long recordings of repeated work are automatically segmented into individual task cycles, so one shift of footage becomes hundreds of training episodes.
Every episode is human-reviewed before delivery, and off-task segments (breaks, interruptions, camera fumbles) are flagged by automatic curation so they never contaminate a policy's training set.
How it works
- 1
Capture
Workers wear head-mounted or body cameras and do their normal tasks; voice notes record intent as it happens.
- 2
Episode pipeline
A 7-model GPU stack annotates every second: actions, objects, pixel-accurate masks, transcripts, and 7-D hand/end-effector state.
- 3
Segment & review
Long recordings auto-split into task cycles; humans review every episode and curation drops off-task segments.
- 4
Robot-format export
Episodes export as LeRobot v3 and RLDS / Open X-Embodiment, with train/valid/test splits.
What you get
- Egocentric capture with per-second annotation
- 7-D hand / end-effector state per frame
- Automatic segmentation of long footage into task episodes
- Human review of every episode + off-task curation
- LeRobot v3 and RLDS / Open X-Embodiment exports
Episodes export as LeRobot v3 and RLDS / Open X-Embodiment style datasets - the formats robotics foundation models and imitation-learning frameworks load natively - plus Parquet and JSONL when you need custom pipelines.
Frequently asked questions
Imitation learning (behavior cloning) trains a robot policy from demonstrations: sequences that pair what was observed at each timestep with the action that was taken. Payana produces this from human work - egocentric video paired with per-frame hand/end-effector state - so robots can learn manipulation from people rather than from expensive robot teleoperation.
Have footage to annotate?
Tell us your task and target format - we handle capture, annotation, review and delivery end to end.
Book an intro call