Skip to main content

Dexset

Case Study: How We Scaled Humanoid Robot Training Data for a VLA Model

Enoch Pakanati

Autonomous driving already ran the experiment humanoid robotics is now repeating. A decade ago, perception teams trained on one city’s dashcam footage, aced the local test loop, and watched the model fall apart one state over; the durable fix was never a cleverer architecture, it was wider data. The client in this case study, a humanoid foundation model team we will keep anonymous, arrived at the same cliff in a different vehicle: task success rising steeply over the first 40 hours of teleoperation data, then flattening in the low 40s (percent) no matter how many more hours they added. Four robots, twelve manipulation tasks, a VLA architecture adapted from the OpenVLA lineage, and a data budget that teleoperation alone would exhaust in a quarter.

The plateau was not a model problem. Their episodes were clean but narrow: one lab, one lighting condition, a dozen object instances. The policy had memorized their room. Held-out object arrangements exposed it immediately, which is the standard failure signature of a diverse-data deficit, the same signature that motivated cross-source efforts like Open X-Embodiment (arxiv.org/abs/2310.08864).

The thesis this program put to the test: generalization is bought with cheap, diverse, QA-gated human hours, and precision with a thin layer of native teleoperation, in that order. This post documents what we actually did over 16 weeks: the corpus design, the retargeting QA that made cheap human data trainable, the costs line by line, and where it worked and where it did not. We publish this level of detail because most vendors will not, and because you should demand it from anyone selling you data, including us.

Key Takeaways – Starting point: 60 teleop hours, 12 tasks, success plateaued in the low 40s on held-out arrangements. – Intervention: 4,000 hours of task-matched human egocentric capture for pretraining plus 300 targeted teleop hours for post-training, delivered over 16 weeks. – Retargeting QA rejected 11 percent of raw egocentric episodes on first pass; every rejected batch was re-solved or re-collected before delivery. – Result: held-out task success moved from the low 40s into the 70s, with the largest gains on tasks where human data added object and layout diversity. – Total data spend was under a quarter of the equivalent teleop-only corpus.

The Starting Point

A data plateau is the point where additional hours of the same distribution stop improving policy performance, and it is diagnosed on held-out conditions, not training tasks. The client’s numbers made the diagnosis easy: in-distribution success near 80 percent, held-out object arrangements in the low 40s, and the gap widening as tasks got more contact-rich.

Their fleet math explained why they could not collect their way out. Four humanoids at 2 to 3 usable teleop hours per robot-day yields roughly 50 hours a week under ideal conditions. Covering the visual and physical diversity they needed, hundreds of object instances across dozens of layouts, would have taken more than a year of fleet time and roughly $1M at our benchmark teleop rates of $150 to $400 per hour. They needed a wider distribution at a price that scaled.

The Corpus Design

The plan followed the data pyramid: a wide base of retargeted human data for pretraining, a thin apex of native teleop for post-training. Three design decisions did most of the work:

  • Task-matched capture, not generic capture. We mirrored their twelve tasks in our capture facilities: same object categories, comparable workspace heights, matched container types. Generic kitchen footage is cheap; footage that matches your deployment distribution is what transfers.
  • Deliberate diversity quotas. Per task: minimum 40 object instances, 8 layouts, 3 lighting conditions, and at least 25 capture workers to vary hand size and motion style. Diversity was tracked as a delivery metric, not left to chance.
  • Teleop reserved for the worst six tasks. The post-training budget went only to tasks still underperforming after pretraining, all of them contact-rich: bimanual tote packing, articulated-container opening, deformable handling.

Delivery formats were LeRobot-compatible datasets (github.com/huggingface/lerobot) with synchronized stereo egocentric video, reconstructed hand pose, and retargeted joint trajectories solved against the exact robot description file pinned in the contract, so their acceptance checks and our delivery gates could never disagree about joint limits.

The QA Layer That Made It Work

Retargeting QA is the set of automated gates that verify human motion remains physically valid after mapping to the robot’s skeleton, and it is the difference between a cheap corpus and a useless one. Human egocentric data has no native robot actions; the labels are manufactured by retargeting, so the labels are only as good as the checks behind them.

Every episode in this program passed five gates before delivery:

  • Joint-limit violation rate below 0.5 percent of frames
  • End-effector position error under 2 cm against reconstructed human hand pose
  • Foot-skate detection: stance-foot translation above 1 cm per frame flags the episode
  • Contact consistency: grasp events must survive retargeting with object-relative pose intact
  • Stream synchronization within one frame at 30 fps

Across the program, 11 percent of raw episodes failed at least one gate on first pass. Most failures were re-solved with adjusted retargeting weights; about a third were re-collected. The client saw every QA report. That transparency clause was in the contract, and we think it should be in yours regardless of vendor.

Costs and Timeline

PhaseWeeksVolumeRate (benchmark)Cost
Capture spec + pilot batch1 to 3120 hrs egocentric$30/hr$3,600
Volume egocentric capture3 to 143,880 hrs$22 to $35/hr~$108,000
Retargeting + QArollingall episodesincluded in ratesincluded
Targeted teleop post-training data9 to 16300 hrs$220/hr avg$66,000
Total164,300 hrs~$178,000

The teleop-only equivalent, 4,300 hours at a midpoint $250 per hour, would have priced at roughly $1.08M before fleet capex and would not have fit inside a year of their robot time. The blended corpus landed at about 16 percent of that figure.

Results, Including the Honest Parts

After pretraining on the retargeted egocentric corpus and post-training on the targeted teleop hours, the client reported held-out task success in the 70s, up from the low 40s. Gains were largest exactly where diversity was the deficit: object-generalization tasks improved by 30-plus points. Gains were smallest on the two highest-precision tasks, sub-centimeter insertion among them, where human video’s missing force signal showed; those tasks improved mainly from the teleop layer, not the pretraining layer.

Two lessons we carry forward. First, human data buys generalization, teleop buys precision, and budgeting means deciding how much of each deficit you have; the framework is laid out in The Complete Guide to Humanoid Robot Training Data. Second, QA rejection rates are a feature. A vendor reporting zero rejections is not running gates.

What We Would Do Differently

Honest retrospectives belong in case studies, so here is ours. We would start the targeted teleop collection two weeks earlier: waiting for pretraining metrics before booking robot time cost us schedule, because the six weakest tasks were predictable from the task list alone. Any task combining deformables with sub-centimeter tolerance was going to need native data; we did not need a training curve to tell us that.

We would also add tactile capture to the teleop layer from day one. The two insertion tasks that improved least were exactly the ones where fingertip force carries the signal that vision cannot. The client has since added grasp-force logging to their fleet, and early batches suggest that is where the next ten points of success rate live. When they publish, we will link it here.

Next Step

The fix is probably distributional, not architectural. Book a demo and we will show you the same QA reports from this program on sample episodes, or start with the cost model in The Complete Guide to Humanoid Robot Training Data.

Frequently Asked Questions

How much training data does a humanoid VLA model need?

In this program: 60 initial teleop hours plateaued, and 4,000 retargeted egocentric hours plus 300 targeted teleop hours broke the plateau. Public reference points align: Figure’s Helix used roughly 500 teleop hours; pretraining corpora run into the thousands of hours.

Sixteen weeks in this case, from capture spec to final delivery, with volume egocentric capture running at roughly 350 hours per week at peak. Teleop-only collection of the same volume would have exceeded a year on a four-robot fleet.

About $178,000 versus roughly $1.08M for the same 4,300 hours via teleoperation at midpoint benchmark rates, or about 16 percent of the teleop-only price.

In this program, yes: held-out success moved from the low 40s to the 70s (percent), consistent with the research direction of HumanPlus (arxiv.org/abs/2406.10454) and GR00T N1 (arxiv.org/abs/2503.14734). The gains concentrated on generalization; precision tasks still needed native teleop data.

At minimum: joint-limit violation rates, end-effector error bounds after retargeting, foot-skate detection, contact consistency checks, and per-frame synchronization tolerance, with rejection statistics reported to you per batch.

Enoch Pakanati
Written by

Enoch Pakanati

Enoch Pakanati is the strategic architect behind DexSet’s mission to become the undisputed market leader in robotics training data. He oversees the company’s growth strategy, focusing on capturing dominant market share across all data modalities required for modern robotics, including egocentric capture, teleoperation, and simulation-to-real data pipelines.

At DexSet, Enoch is responsible for transforming the company’s deep technical capabilities into a market-leading brand that foundation model labs and robotics OEMs trust implicitly. He focuses on scaling DexSet’s global footprint and ensuring the company stays ahead of the industry’s rapidly evolving data needs. His leadership is centered on one objective: making DexSet the singular, global standard for the data that powers the robotics revolution.