Why Data Quality & QA for Robotics Datasets Is the Biggest Bottleneck in Physical AI
A question keeps resurfacing in r/robotics threads and on our buyer calls, phrased a dozen ways but always the same underneath: we added 500 hours of demonstrations, so why did the policy get worse? The asker almost always suspects a modeling bug. In our experience the compute is usually fine and the architecture is usually fine; the bottleneck is that a meaningful slice of the new data was quietly hostile to training. That is the argument of this post, stated plainly: usable hours, not raw hours and not compute, are the binding constraint on physical AI right now.
The confusion persists because robot data has failure modes that no amount of scrolling through thumbnails will reveal. A 35 millisecond lag between camera and control streams looks like nothing to a human and looks like a teacher who acts before seeing to a policy. The industry inherited its instincts from vision and language, where a noisy sample costs one gradient step. In closed-loop imitation learning, a noisy demonstration costs behavior.
This post makes that case with three kinds of evidence: the mechanics of how bad demos poison policies, the numbers on how much collected data actually fails inspection, and the metric that should replace raw hours in every planning meeting.
At DexSet we run collection and QA for VLA and humanoid teams, and we reject 10 to 30 percent of our own raw hours before delivery. Those rejection logs are where most of what follows comes from.
Key Takeaways – In our production pipelines, 10 to 30 percent of raw teleoperation hours fail QA. Raw hours are a vanity metric. – Imitation learning compounds errors: bad demos push policies into unfamiliar states where they fail further (the DAgger argument, Ross et al.). – The most common silent killers are action-observation desync, camera dropouts, and wrong success labels. – Usable-hour yield is the metric that predicts training outcomes. Track it per rig, per operator, per vendor. – QA overhead of $5-$12 per hour is cheap against collection costs of $28-$60 per hour and wasted GPU runs.
The Bottleneck Is Usable Hours, Not Hours
The real unit of progress in robot learning is the usable hour: a demonstration hour that survives signal, semantic, and distribution checks. Teams plan in raw hours because that is what invoices and dataset cards report, but the two numbers diverge by 10 to 30 percent in our production experience, and the gap is invisible until training exposes it.
Consider what that gap does to a budget. At $40 per collected hour, a 1,000-hour order costs $40,000. At 75 percent yield, you paid $53 per usable hour without knowing it, and worse, the 250 failing hours do not just waste money. Some fraction of them, if they reach training, subtract performance. The bottleneck compounds: you pay for data that hurts, then pay GPU time to discover it hurt, then pay again to recollect.
Open X-Embodiment gives the field-scale version of this story. Pooling 22 embodiments into one corpus required heavy normalization of action spaces, camera conventions, and labels before the data was jointly usable (arXiv:2310.08864). The lesson generalizes down to a single lab: heterogeneous, unaudited data is not a dataset yet. It is raw material.
How Bad Demos Poison Policies
A corrupted demonstration damages an imitation policy through compounding errors and distribution shift, not through simple label noise. Behavior cloning trains on expert states, but at deployment the policy visits its own states. Ross, Gordon, and Bagnell showed that small per-step errors accumulate quadratically in the horizon under this mismatch, which is the core motivation for DAgger (arXiv:1011.0686). Every bad demo widens the per-step error, and the trajectory dynamics amplify it from there.
The failure modes are specific:
- Action-observation desync teaches the policy that action at time t follows the observation at t minus two frames. It learns a systematically wrong causal mapping, then acts early or late at deployment.
- Mislabeled successes put failed grasps inside the “imitate this” set. The policy reproduces the failure with confidence.
- Operator recovery segments left unsegmented teach oscillation: approach, miss, retreat, retry becomes a valid strategy in the policy’s eyes.
- Black frames and dropouts train the vision encoder to tolerate, and sometimes act on, absent information.
The ACT authors were explicit that fine manipulation demands coping with noisy, stochastic human demonstrations, and their action-chunking design exists partly to absorb that noise (arXiv:2304.13705). Architecture can absorb noise. It cannot absorb systematically wrong data. Our internal ablations back this up: cleaned subsets beat larger raw sets on held-out task success, most dramatically on long-horizon tasks.
What Actually Fails, and How Often
The distribution of QA failures is stable enough across our fleet that you can budget against it. Here is what our rejection logs look like on a typical mixed teleoperation workload:
| Failure category | Share of rejected hours (our fleet) | Detectable by automation? |
|---|---|---|
| Camera dropout / black frames | 20-30% | Yes |
| Action-observation desync >1 frame | 15-25% | Yes |
| Timestamp non-monotonicity | 10-15% | Yes |
| Gripper state mismatch | 10-15% | Yes |
| Episode length outliers | 10-15% | Yes |
| Wrong success labels, sloppy strategy | 15-25% | No, needs human review |
Two things worth noticing. First, most failures are automatable catches, which is why cheap Tier 1 checks recover most of the yield gap. Second, the human-only category is the most dangerous per hour, because success labels gate what enters training in most pipelines. We audit success labeling accuracy on every batch and hold the line at 95 percent; details and thresholds are in our full guide to data quality and QA for robotics datasets.
Why the Bottleneck Is Getting Tighter, Not Looser
The quality bottleneck tightens as models scale, because larger training corpora mix more sources and longer-horizon tasks amplify compounding errors. A team fine-tuning a VLA like OpenVLA or π0 on task-specific demonstrations is pooling its own data with pretraining distributions it never audited (arXiv:2406.09246). Any convention mismatch between the fine-tuning data and the pretraining recipe, action normalization, camera framing, gripper semantics, lands directly on the model.
Task horizon makes it worse. A 5-second pick task gives a desynced policy little room to drift. A 90-second bimanual assembly gives per-step errors ninety seconds to compound, which is why our rejection thresholds tighten, not loosen, for long-horizon collections. The tasks the industry is moving toward, mobile manipulation, multi-stage kitchen and logistics work, are precisely the tasks least forgiving of quality debt. Teams that treat QA as a phase-two problem are scheduling that debt to come due at their most expensive moment.
The Metric That Should Run Your Planning Meetings
Usable-hour yield, the fraction of collected hours surviving QA, is the single most decision-relevant number in a robot data operation. It converts vendor pricing into real pricing. It exposes rig degradation early. It turns “we need more data” into “we need yield above 85 percent before we scale collection.”
Getting it requires actually running QA, which costs $5 to $12 per hour in our benchmarks across automated checks, human spot review, and policy-in-the-loop validation. Against $28 to $60 per hour collection and five-figure GPU runs, that overhead pays for itself the first time it stops one poisoned batch from reaching training. If you buy data rather than collect it, demand the yield number and the rejection criteria in the contract. A vendor who cannot state their own rejection rate has never measured it, and a vendor who claims zero rejections is telling you the QA process, not the data, is the thing that does not exist.
Next Step
The full three-tier pipeline, the automated check catalogue with thresholds, and the cost breakdown live in our complete guide: Data Quality & QA for Robotics Datasets. If you want your own yield number, book a pipeline review and we will audit a sample of your corpus.
Frequently Asked Questions
Why is data quality the main bottleneck in physical AI rather than compute?
Because bad demonstrations subtract performance rather than merely wasting compute. Compounding errors in imitation learning mean a 15 percent bad-data fraction can dominate training outcomes, and no additional compute corrects a systematically wrong causal mapping.
How much collected robot data typically fails QA?
Our production pipelines reject 10 to 30 percent of raw teleoperation hours. New rigs and long-horizon bimanual tasks sit at the high end; mature rigs on short tasks sit near 10 percent.
What is usable-hour yield?
The fraction of collected hours that pass all QA tiers and are fit for training. It is the number that converts nominal per-hour pricing into true per-usable-hour cost.
Can automated checks alone solve robot data quality?
No. Automation catches signal-integrity failures like desync and dropouts, which are most rejections by volume, but success-label errors and distribution problems need human review and policy-in-the-loop validation.
What does robotics data QA cost?
In our benchmarks, $5 to $12 per collected hour on top of collection costs of $28 to $60 per hour, or roughly 10 to 30 percent overhead.
Enoch Pakanati
Enoch Pakanati is the strategic architect behind DexSet’s mission to become the undisputed market leader in robotics training data. He oversees the company’s growth strategy, focusing on capturing dominant market share across all data modalities required for modern robotics, including egocentric capture, teleoperation, and simulation-to-real data pipelines.
At DexSet, Enoch is responsible for transforming the company’s deep technical capabilities into a market-leading brand that foundation model labs and robotics OEMs trust implicitly. He focuses on scaling DexSet’s global footprint and ensuring the company stays ahead of the industry’s rapidly evolving data needs. His leadership is centered on one objective: making DexSet the singular, global standard for the data that powers the robotics revolution.