A managed vision intelligence service by Baaz

Real Human Work,
Robot-Ready Datasets

We capture your real-world tasks on wearable cameras, annotate every moment, and deliver structured, training-ready AI & robotics datasets - end to end, as a managed service.

Hand detection accuracy
94%

Hand detection accuracy

Action detection accuracy
92%

Action detection accuracy

Object detection accuracy
88%

Object detection accuracy

Training formats
7

Training formats

Exports to every major training format

LeRobotRLDSOpen X-EmbodimentDROIDAgiBot WorldParquetJSONL
The Service

Everything, Handled

From wearable capture to production-ready datasets - delivered as a managed service by Baaz. Every step automated, every label verified, every export ready for training with always on human in the loop.

Wearable Capture

Record real work exactly as it happens with smart glasses, helmet cams, or mobile devices. First-person (egocentric) footage is the richest signal for robot learning - and we handle the whole capture setup.

  • Offline-capable, resumable field uploads
  • Voice intent recorded alongside video
  • Live streaming for remote supervision

AI Annotation Pipeline

A 7-model GPU stack understands every second of footage: a vision-language model narrates the action, detection models find every object and tool, segmentation traces pixel-accurate masks, and speech becomes searchable transcripts.

  • Action + scene understanding per frame
  • Open-vocabulary detection - not limited to fixed classes
  • Hand tracking → 7-D end-effector state for robotics

Rich, Structured Labels

Every second carries objects with bounding boxes and masks, the action being performed, a scene description, the spoken transcript, and confidence scores - structured data your training code consumes directly.

  • Objects · actions · transcripts · confidence
  • Off-task seconds auto-flagged by curation
  • Everything searchable across your library

Annotation Workspace

A full labeling studio for images and video frames: bounding boxes, polygons, semantic masks, keypoints and skeletons, and classification - with one-click AI smart polygons doing the tedious tracing.

  • Click an object → AI finds the exact boundary
  • Project types keep every dataset format-valid
  • Keyboard-first review: label → save → next

Human Review & Curation

AI proposes, people verify. Every episode passes a review workspace with annotated playback, per-second quality timelines, and keep/exclude controls - so only clean, on-task data reaches your datasets.

  • Frame-accurate annotated playback
  • Reviewer overrides tracked & auditable
  • Success/failure outcome labeling

Episode Intelligence

Individual seconds merge into complete task episodes: step-by-step workflows, the tools and objects involved, and knowledge graphs linking actions across your whole footage library. Long recordings split into task cycles automatically.

  • Automatic task segmentation of long recordings
  • Workflow steps with timing & tooling
  • Cross-episode knowledge graph

Versioned Dataset Delivery

Datasets ship as immutable versions with train/valid/test splits, quality metrics, and one-click downloads in every major format - for computer vision and for robot learning.

  • COCO · YOLO · Pascal VOC · PNG masks
  • LeRobot v3 · RLDS · Parquet · JSONL
  • Reproducible versions, secure delivery

Model Training

We don't stop at datasets. Train a production-ready detection, segmentation, keypoint or classification model on your labeled data - on our GPUs, with live progress - and download deployable model weights the same day.

  • Train on Payana labels, episode data, or your own uploaded dataset
  • Retrain anytime with more data - or fine-tune your existing model
  • Honest accuracy metrics reported with every trained model

Private by Design

Your footage and datasets live on dedicated, private cloud infrastructure - isolated storage, locked-down access, and time-limited download links. Your data trains your models, nobody else's.

  • Private, isolated storage - never a public pool
  • Isolated processing on dedicated infrastructure
  • Never used to train third-party models
Workflow

How It Works

Data collection, annotation, and delivery - take the full pipeline, or just the piece you need. We collect for you, annotate your own data, or both.

  1. Data collection

    Two ways in: send us the images and video you already have, or we collect for you - wearable and egocentric capture, field recording, and frames extracted from your footage at whatever rate the project needs.

    You provide itOr we capture itImages · video · video → frames
  2. Annotation

    human in the loop

    Turning raw pixels into machine-readable labels: bounding boxes for detection, pixel-accurate polygons and masks for segmentation, keypoints and skeletons for pose, classes and persistent tracking IDs across video. Every label is drawn and verified by our annotation team before it counts.

    Boxes · polygons · masksKeypoints · tracking IDsEvery label verified
  3. Delivery

    Your labeled data ships as a versioned dataset with train/valid/test splits and quality checks - in the format you train with, ready for your pipeline the day it lands.

    COCOYOLOPascal VOCPNG masksLeRobotRLDS
Annotation workspaceplants · frame 0042
REC · headmount00:42
01 · Head-mounted camera records the work
See it done

From Camera to Robot, In One Take

Recording the work, first person
The Output

What Labeled Data Looks Like

Every image and frame comes back annotated - bounding boxes, pixel-accurate polygons, and keypoints - exactly as your training code needs them.

Object Detection
Object detection annotation example: pedestrians, bags and poles labeled with bounding boxes

Bounding boxes with class labels and confidence scores - exported as COCO JSON, YOLO TXT or Pascal VOC.

Instance Segmentation
Instance segmentation annotation example: warehouse crates, pallets and roof tiles outlined with pixel masks

AI smart polygons trace exact object boundaries - one click instead of twenty vertices, exported as COCO segmentation or PNG masks.

Keypoints
Keypoint annotation example: runners tracked with full-body pose skeletons

Skeleton keypoints track hands and tools through every frame - the signal that turns human work into robot-trainable actions.

Why Payana

One Service, Humans at the Core

Self-serve tools leave the labeling to your team. Traditional vendors stop at raw annotations. Payana puts trained annotators on your data from collection to delivery - every label drawn and verified by people, then shipped as versioned datasets in every major training format.

Data collection done for you (wearable / field capture)

Egocentric and field capture when you don't have data yet

PayanaYes
Self-serve toolsNo
Labeling vendorsNo
Data collection vendorsYes

Labeling done for you

Our team + AI pipeline annotate; you review and receive datasets

PayanaYes
Self-serve toolsNo
Labeling vendorsYes
Data collection vendorsNo

AI-assisted annotation (one-click smart polygon)

One click finds the exact object boundary - on our own GPUs

PayanaYes
Self-serve toolsPartial
Labeling vendorsNo
Data collection vendorsNo

Boxes, polygons, segmentation, keypoints, classification

PayanaYes
Self-serve toolsYes
Labeling vendorsYes
Data collection vendorsNo

Video → sampled frames for image labeling

PayanaYes
Self-serve toolsYes
Labeling vendorsNo
Data collection vendorsNo

Automatic annotation of wearable / egocentric footage

7-model GPU pipeline: vision-language model, detection, pixel-accurate masks, speech, hand pose

PayanaYes
Self-serve toolsNo
Labeling vendorsYes
Data collection vendorsNo

CV exports: COCO JSON · YOLO · Pascal VOC · PNG masks

PayanaYes
Self-serve toolsYes
Labeling vendorsPartial
Data collection vendorsNo

Robot-learning exports: LeRobot · RLDS · Open X-Embodiment

PayanaYes
Self-serve toolsNo
Labeling vendorsNo
Data collection vendorsNo

Model training - labeled data to deployable model weights

Trained on our GPUs from your labels or your own dataset; retrain or fine-tune with more data anytime

PayanaYes
Self-serve toolsPartial
Labeling vendorsNo
Data collection vendorsNo

Dataset versioning with train/valid/test splits

PayanaYes
Self-serve toolsYes
Labeling vendorsNo
Data collection vendorsNo

Human review & curation included in the price

PayanaYes
Self-serve toolsNo
Labeling vendorsYes
Data collection vendorsPartial

Your data on dedicated private infrastructure

Private, isolated infrastructure; never pooled, never used to train third-party models

PayanaYes
Self-serve toolsPartial
Labeling vendorsPartial
Data collection vendorsPartial

Frequently Asked Questions

Payana is the managed vision intelligence service by Baaz. We turn raw images, videos, and wearable-camera footage into training-ready datasets: AI-assisted labeling, human review, dataset versioning with train/valid/test splits, and export to COCO JSON, YOLO, Pascal VOC, PNG masks, LeRobot, RLDS, Parquet and JSONL.

Connect with us

Book an intro call

Pick a slot that works for you - we'll walk through your use case, the capture setup, and the dataset formats you need.

July 2026

Sun
Mon
Tue
Wed
Thu
Fri
Sat

Wed 22

No open slots on this date.

Start your dataset

Ready for Better Training Data?

Tell us the tasks and the format you need - the Baaz team captures, annotates, and delivers your dataset end to end.