Skip to main content

Dexset

Comparing Humanoid Robot Training Data Approaches: Pros, Cons & Costs

Sainath Gupta

You are six weeks from a funded milestone demo, your one humanoid logs two usable teleop hours on a good day, and the task list your investors saw has forty entries. In the planning meeting someone says the obvious thing: “we’ll collect it all by teleoperation, that’s what Figure did.” Sometimes that is right. Often it commits a seed-stage budget to $250-per-hour data for tasks that $25-per-hour human capture would have covered, or the reverse, a pretraining corpus full of retargeted human video when what the policy actually needed was 200 clean hours on the target robot.

The trap exists because the capture methods look interchangeable on a slide. They all produce “episodes.” Underneath, they differ by an order of magnitude in cost, by a week versus a quarter in scaling speed, and by whether the actions in the file are native to your robot or borrowed from a human body and translated.

Our thesis: the right capture method is a function of training stage and task physics, never of preference or precedent, and the decision can be made from four measurable axes. This post compares the five approaches on those axes: cost per usable hour, throughput, action fidelity, and failure modes. The numbers are DexSet’s own operating benchmarks; we run four of the five methods daily and rent GPUs for the fifth like everyone else.

Key Takeaways – Five approaches: human egocentric video, mocap suits, VR teleoperation, exoskeleton capture, and simulation. No single one covers a full humanoid program. – Cost per usable hour spans two orders of magnitude: under $5 (sim) to $400 (on-robot teleop). – The real differentiator is action fidelity: native robot actions (teleop) vs retargeted human motion (egocentric, mocap) vs synthetic physics (sim). – Retargeting error is the hidden cost of the cheap methods; QA gates decide whether human data is a bargain or a write-off. – Match method to training stage: human data for pretraining, sim for locomotion and augmentation, teleop for post-training.

The Five Approaches Defined

A humanoid data approach is the combination of capture hardware, human role, and label type used to produce training episodes. The five in production use today:

  • Human egocentric video: head-mounted cameras (Project Aria, Quest 3 passthrough, custom stereo) record a person doing tasks with their own body. No robot involved. Actions must be reconstructed and retargeted.
  • Mocap suit capture: IMU body suits record full-body human kinematics directly, Xsens-style, at 60 to 240 Hz. Clean joint trajectories, no cameras required, still human-embodied.
  • VR teleoperation: an operator pilots the actual humanoid through a headset and tracked controllers or body tracking. Observations and actions are native to the robot. This is the method behind Figure’s Helix corpus (roughly 500 hours).
  • Exoskeleton / leader-arm capture: the human moves inside a passive structure kinematically matched to the robot’s arms, so joint-space actions are recorded without inverse-kinematics estimation. ALOHA’s leader arms (arxiv.org/abs/2304.13705) are the canonical low-cost version.
  • Simulation: synthetic episodes from Isaac Lab or MuJoCo with domain randomization over dynamics, textures, and terrain. The default for locomotion RL on platforms like Unitree G1/H1 class hardware.

Master Comparison Table

The table below is the one we wish every RFP included. Ranges are DexSet operating benchmarks (first-hand), 2026.

ApproachCost per usable hourHardware capexThroughput per station-dayAction fidelityHand/finger detailScales by
Human egocentric video$15 to $40$500 to $6,0005 to 6 hrsRetargeted (lowest)Video-reconstructedHiring (days)
Mocap suit$40 to $90$10,000 to $45,0004 to 5 hrsHuman joint spaceWeak without glovesSuits + hiring (weeks)
Exoskeleton capture$80 to $200$20,000 to $60,0003 to 4 hrsMatched joint spaceGood (upper body)Stations (months)
VR teleoperation$150 to $400$16,000 to $200,000+2 to 3 hrsNative robot actionsBest availableRobot fleet (months)
Simulation$0.50 to $5 (hr-equiv)Compute onlyUnlimitedSynthetic physicsModeled, imperfectGPU budget (hours)

Two readings of this table matter more than any single cell. First, cost and fidelity are inversely arranged: the cheapest data is furthest from your robot’s embodiment. Second, the “scales by” column is the one founders skip: teleop scales with robot procurement lead times, egocentric capture scales with hiring.

Pros and Cons by Approach

Human Egocentric Video

Egocentric capture wins on price, diversity, and scale, which is why the human-data hypothesis (humans as the largest humanoid dataset, per Ego4D, EgoExo4D, and the GR00T N1 pyramid) keeps gaining ground, reportedly including inside Tesla’s Optimus program. Pros: $15 to $40 per hour, unlimited scene variety, 5 to 6 usable hours per capture-day. Cons: no robot action labels, no force signal, and everything depends on retargeting quality. If your vendor cannot show you their retargeting QA gates, the discount is fictional.

Mocap Suits

Mocap suits win on kinematic cleanliness for whole-body motion: gait, turning, load carrying, crouching. Retargeting pipelines like H2O (arxiv.org/abs/2403.04436) and HumanPlus (arxiv.org/abs/2406.10454) consume exactly this format. Pros: occlusion-free joint trajectories, high frame rates. Cons: IMU drift on long takes, poor fingertip detail, and suits cost real money. We pair suits with hand-tracking gloves or egocentric video on manipulation tasks.

VR Teleoperation

Teleop wins on truth. Observations come from the robot’s cameras, actions are the robot’s own joint commands, and contact happens with the robot’s real hands. Pros: zero embodiment gap, the gold standard for post-training. Cons: $150 to $400 per delivered hour, 2 to 3 usable hours per robot-day, operator skill curves (about two weeks on our stations before pass rates stabilize), and fleet lead times.

Exoskeleton Capture

Exoskeleton rigs are the compromise position: direct joint-space actions without occupying a robot. Pros: strong for dexterous upper-body tasks, no IK guesswork. Cons: workspace constraints, task variety limits, and station costs that put it closer to teleop than to egocentric capture.

Simulation

Sim wins wherever physics is honest. Locomotion policies trained with domain randomization in Isaac Lab transfer to real hardware routinely; this is why bipedal walking is no longer the hard part. Pros: near-zero marginal cost, perfect labels, infinite parallelism. Cons: contact-rich manipulation, deformables, and cloth still cross the sim-to-real gap badly, so manipulation data must come from the physical world.

Decision Matrix: Which Approach When

The right approach depends on training stage and task family, not on preference:

Your situationRecommended mix
Pretraining a humanoid VLA from scratch70 to 85 percent egocentric + mocap, 10 to 20 percent sim augmentation, 5 to 10 percent teleop
Locomotion and whole-body controlSim-first (Isaac Lab), mocap for motion style, minimal real capture
Dexterous manipulation post-trainingTeleop and exoskeleton hours concentrated on failing tasks
Budget under $100k, need signal fastEgocentric capture matched to your task list, plus sim; defer teleop
Robot fleet already deployedTeleop during idle time, human data to widen task coverage

The blended pyramid is not a compromise; it is the shape the evidence supports, from Open X-Embodiment’s cross-source lesson (arxiv.org/abs/2310.08864) to GR00T N1’s explicit data pyramid (arxiv.org/abs/2503.14734). The full economics, including a worked 10,000-hour budget, are in our pillar: The Complete Guide to Humanoid Robot Training Data.

Next Step

If you are choosing between these approaches right now, do not decide from a table alone, including this one. Book a sample review and we will send matched episodes from our egocentric, mocap, and teleop pipelines with QA reports attached, so you can evaluate the actual files against your training stack. The full method guide is here: The Complete Guide to Humanoid Robot Training Data.

Frequently Asked Questions

What is the cheapest way to collect humanoid robot training data?

Human egocentric video, at $15 to $40 per usable hour in our benchmarks, followed by simulation at under $5 per hour-equivalent. Egocentric data requires retargeting and QA before it is trainable; sim covers locomotion but not contact-rich manipulation.

For post-training, yes: teleop is the only method producing native robot observations and actions. For pretraining volume, usually not; retargeted human data delivers diversity at a fraction of the price, which is why programs reserve teleop for final-mile tuning.

Mocap suits for whole-body motion quality (gait, balance, carrying), egocentric video for scale, scene diversity, and vision-paired manipulation. Many programs use both: suits for locomotion style, glasses for task coverage.

Current physics engines model rigid-body dynamics well but struggle with fingertip friction, deformable objects, cloth, and liquids. Policies trained purely in sim lose success rate on contact-rich tasks after transfer, so real demonstrations remain necessary.

Ask for the QA gates applied per episode: synchronization tolerance, joint-limit violation rates after retargeting, end-effector error bounds, foot-skate detection, and the re-collection policy for rejected batches. Our RFP Scorecard in the pillar guide lists all ten dimensions.

Sainath Gupta
Written by

Sainath Gupta

Sainath Gupta is the visionary leader of DexSet, a company dedicated to building the foundational data layer for the robotics revolution. Sainath recognized early on that the primary bottleneck for physical AI wasn't hardware or model architecture, but the lack of high-quality, scalable training data.

At DexSet, he has pioneered a "manufacturing-first" approach to data collection. This includes the development of transparent cost-per-hour benchmarks and the scaling of global networks for egocentric and teleoperation capture. Sainath is a vocal advocate for the industry’s shift toward treating human experience as the primary pretraining substrate for robots. His leadership at DexSet is focused on one goal: providing the millions of hours of high-fidelity data required to bring humanoid robots out of the lab and into the real world.