Why Humanoid Robot Training Data Is the Biggest Bottleneck in Physical AI
More on-robot data is the wrong default for a humanoid program, and the best-funded labs are the ones proving it. If teleoperation hours were the answer, the teams with the deepest pockets would simply buy more of them; instead, NVIDIA architected GR00T N1 around a data pyramid whose wide base is human video (arxiv.org/abs/2503.14734), and Optimus data collection has reportedly shifted toward human-worn capture rigs. Meanwhile the formerly hard parts got easy: a Unitree G1 lists near $16,000 (unitree.com), and OpenVLA weights are a download on Hugging Face. The one ingredient you cannot download or buy off a pallet is ten thousand hours of good demonstrations.
The gap exists for a structural reason. Language models trained on a web that humanity spent thirty years writing for free. There is no internet of robot actions. Every hour of humanoid experience has to be manufactured, one physical hour at a time, by a person, a robot, or a simulator, and two of those three are expensive.
Our thesis: data supply, not modeling, now sets the pace of every humanoid program, and the escape is manufactured human experience, not more robot time. This post argues that claim three ways: the throughput arithmetic of teleoperation, the priced-out escape routes, and the scaling pattern we see working across the programs we supply. The numbers come from DexSet’s own capture operations: egocentric rigs, teleop stations, and the QA pipeline behind them.
Key Takeaways – LLMs had the web; humanoids have nothing comparable. Robot experience must be manufactured hour by hour. – Figure reported roughly 500 hours of teleoperation behind Helix. Language models pretrain on web-scale corpora measured in trillions of tokens. That is the scale of the gap. – On-robot teleoperation runs $150 to $400 per delivered hour in our benchmarks and scales with robot fleet size, which has months-long lead times. – Human egocentric capture runs $15 to $40 per hour and scales with hiring, which takes days. This is why humans are becoming the primary humanoid dataset. – The teams escaping the bottleneck use a pyramid: human video for pretraining, simulation for augmentation, teleop only for final-mile post-training.
Why the Data Bottleneck Exists
The humanoid data bottleneck is the gap between the hours of experience modern robot learning methods need and the rate at which real humanoid hours can be produced. It exists because humanoid data generation is serial, embodied, and supervised: one robot, one operator, one wall-clock hour per data hour, minus resets and failures.
Consider the arithmetic. In our teleop operations, a humanoid station yields 2 to 3 usable hours per robot-day after setup, resets, retries, and QA rejection. A ten-robot fleet running six days a week produces roughly 150 usable hours weekly. At that rate a 10,000-hour corpus takes about 16 months, and that assumes no downtime, which anyone who has maintained a fleet of humanoids will find optimistic. Actuators fail. Hands fail more.
Contrast the model side. VLA architectures are public (OpenVLA, arxiv.org/abs/2406.09246; pi-0, arxiv.org/abs/2410.24164; GR00T N1, arxiv.org/abs/2503.14734). Compute is rentable by the hour. Data alone is neither public nor rentable at humanoid scale, which is why it sets the pace of the whole program.
Why Humanoids Have It Worse Than Arms
Humanoid data is harder to collect than manipulator data because whole-body embodiment multiplies degrees of freedom, safety overhead, and failure modes. A tabletop ALOHA cell needs two leader arms and a folding table. A humanoid needs balance control running under everything, a gantry or spotter, and an operator interface that maps a human body to 23-plus actuated joints.
Three multipliers stack up:
- Dimensionality. More joints means more channels to synchronize, more ways for retargeting to fail, and more data needed to cover the state space.
- Two data regimes, not one. Humanoids need whole-body control data (locomotion, balance) and manipulation data (grasping, tool use). Sim covers the first well; the second still demands real demonstrations. See the capture approaches comparison for the full breakdown.
- Hands. Dexterous hands are the most fragile, most expensive, and most data-hungry subsystem. Teleoperating them well is a skill; a new operator on our stations needs about two weeks before their episodes pass QA at target rates.
The result: cost per usable humanoid hour lands an order of magnitude above cost per arm hour, at exactly the moment the field decided it needs orders of magnitude more hours.
The Three Escape Routes, Priced
There are three ways past the bottleneck: simulate, teleoperate faster, or mine human motion. The table uses DexSet benchmark ranges from our own operations.
| Escape route | Cost per usable hour | Scaling constraint | Where it wins | Where it fails |
|---|---|---|---|---|
| Simulation (Isaac Lab, MuJoCo) | $0.50 to $5 (hour-equivalent) | GPU budget only | Locomotion, balance, augmentation | Contact-rich manipulation, deformables |
| On-robot teleoperation | $150 to $400 | Robot fleet size, operator supply | Embodiment-true post-training | Throughput: 2 to 3 usable hrs per robot-day |
| Human motion capture (egocentric video, mocap suits) | $15 to $90 | Hiring and rig supply | Pretraining scale, task diversity | Needs retargeting plus QA to be usable |
Each route alone is a trap. Sim-only programs ship robots that walk beautifully and fumble a coffee cup. Teleop-only programs run out of money or calendar. Human-data-only programs discover retargeting error the hard way. The pattern that works is the pyramid NVIDIA formalized around GR00T: broad human data at the base, synthetic augmentation in the middle, a thin layer of teleop at the top.
The Human-Data Hypothesis
The human-data hypothesis holds that humans are the largest available source of humanoid training data, because human bodies are kinematically close enough to humanoids for recorded human motion to transfer through retargeting. Seven billion general-purpose “humanoids” already perform every task we want robots to do, every day, for free.
The research base has matured fast. Ego4D (arxiv.org/abs/2110.07058) and EgoExo4D (arxiv.org/abs/2311.18259) proved egocentric capture at thousands-of-hours scale. HumanPlus (arxiv.org/abs/2406.10454) and H2O (arxiv.org/abs/2403.04436) showed human motion retargeting onto real humanoids. Optimus data collection has reportedly shifted toward human-worn capture rigs. The economics explain why: in our operations, a person wearing a capture rig produces 5 to 6 usable hours a day at $15 to $40 per hour, and adding capacity is a hiring problem, not a procurement problem.
The honest caveat: human data arrives without robot action labels, and retargeting is where quality is decided. Our pipeline enforces hard gates per episode: joint-limit violations below 0.5 percent of frames, end-effector error under 2 cm post-retargeting, and foot-skate detection. Cheap data that fails silently is more expensive than teleop. Cheap data with enforced QA is the way out of the bottleneck.
What Scaling Actually Looks Like
A scalable humanoid data program sequences its sources by cost, not by fidelity. The pattern we see succeed:
- Weeks 1 to 4: Define the task list and capture spec. Match human capture scenarios to robot deployment scenarios (same object classes, same layouts).
- Months 1 to 4: Run egocentric and mocap capture at volume for pretraining. Thousands of hours, tens of thousands of dollars per month, not hundreds.
- Continuously: Generate simulated locomotion and augmentation data in Isaac Lab against the same task set.
- Months 3 onward: Spend teleop hours only on the tasks where the pretrained policy underperforms. This is where $250-per-hour data earns its price.
One client program moved held-out manipulation success from the low 40s to the 70s (percent) on this sequence, at under a quarter of teleop-only cost. Details in the VLA case study, and the full method, including modality definitions and the complete cost model, lives in our pillar: The Complete Guide to Humanoid Robot Training Data.
Next Step
If your data plan still assumes teleop-only volume, run the numbers against the pyramid model before you commit the budget. The Complete Guide to Humanoid Robot Training Data includes the full cost tables and a downloadable RFP Scorecard for pressure-testing vendors, including us.
Frequently Asked Questions
Why is training data the bottleneck for humanoid robots?
Because model architectures and hardware are now widely available, while robot experience must still be manufactured one physical hour at a time. Teleoperation yields only 2 to 3 usable hours per robot-day, so data production, not modeling, sets program timelines.
How much does it cost to scale humanoid robot training data?
Benchmark ranges from our operations: $15 to $40 per hour for human egocentric capture, $40 to $90 for mocap, $150 to $400 for on-robot teleoperation, under $5 per simulated hour-equivalent. A 10,000-hour human-first corpus runs roughly one fifth the cost of the same volume in teleop.
Can simulation alone solve the humanoid data problem?
No. Simulation handles locomotion and balance well through domain randomization, but contact-rich manipulation (deformables, fingertip friction, cloth) still transfers poorly, so manipulation policies need real demonstrations.
What is the human-data hypothesis for humanoids?
The hypothesis that recorded human motion, captured through egocentric video and mocap, is the largest viable training corpus for humanoids because human and humanoid kinematics are close enough for retargeting to work, as demonstrated by HumanPlus, H2O, and the GR00T N1 data pyramid.
How many hours did existing humanoid models train on?
Figure reported roughly 500 hours of teleoperation behind Helix. Human-video pretraining corpora in the research literature run into the thousands of hours (Ego4D exceeds 3,600 hours).
Enoch Pakanati
Enoch Pakanati is the strategic architect behind DexSet’s mission to become the undisputed market leader in robotics training data. He oversees the company’s growth strategy, focusing on capturing dominant market share across all data modalities required for modern robotics, including egocentric capture, teleoperation, and simulation-to-real data pipelines.
At DexSet, Enoch is responsible for transforming the company’s deep technical capabilities into a market-leading brand that foundation model labs and robotics OEMs trust implicitly. He focuses on scaling DexSet’s global footprint and ensuring the company stays ahead of the industry’s rapidly evolving data needs. His leadership is centered on one objective: making DexSet the singular, global standard for the data that powers the robotics revolution.