Why Scaling Robot Data Programs Is the Biggest Bottleneck in Physical AI
On a shelf above cell 3 on our capture floor sits a coffee tin full of worn-out gripper fingertips. A pair lasts a few thousand episodes before the silicone goes glossy and grip force starts to drift, so the tin fills week by week, one physical interaction at a time. I keep it there because it is the most honest chart we own of where physical AI’s constraint actually sits: not compute, not architecture, but episodes on disk. Language models trained on trillions of scraped tokens. The largest public robot dataset, Open X-Embodiment, holds just over one million trajectories, and assembling it took 21 institutions pooling years of work (arXiv:2310.08864).
The gap exists because robot data has a property no other modality shares: every sample requires a physical event. A rig has to exist, an operator has to move it, a scene has to be set and reset, and someone has to verify the episode is usable. There is no crawler for the physical world. Money helps but does not remove the constraint; Google’s RT-1 collection ran 13 robots for 17 months to reach roughly 130,000 episodes (arXiv:2212.06817).
The common response is buying more rigs, followed by surprise when output barely moves. That surprise points at this post’s thesis: the bottleneck is rarely hardware. It is the program around the hardware: operators, QA, ingest, and scheduling.
This post breaks down where the hours actually go, why doubling rigs does not double data, and what the throughput math says about fixing it. The numbers come from DexSet’s own capture operations, where we track cost per usable hour weekly across teleop and egocentric cells.
Key Takeaways – Robot data cannot be scraped; every episode is a physical event, which caps scaling at the speed of operations, not compute. – Raw hours are not usable hours. QA rejection, setup, resets, and downtime cut an 8-hour shift to 5.5 to 6.5 usable hours. – Throughput = rigs × shifts × usable hours per shift. 10 rigs × 2 shifts × 6 usable hours = 120 usable hours per day. – The choke point moves as you scale: first rigs, then operators, then QA, then ingest. Fix them in that order or output stalls. – Public reference points (RT-1, DROID, Open X-Embodiment) all reached scale through organization, not lab heroics.
What Makes Robot Data a Bottleneck?
The robot data bottleneck is the structural limit on how fast physically grounded training data can be produced, and it exists because collection speed is bounded by rigs, human operators, and real-world scene resets rather than by compute or bandwidth. Text and image models sidestepped this by consuming data that already existed. Robot actions paired with observations do not exist until someone creates them, one episode at a time.
Consider the arithmetic. A skilled teleoperator produces perhaps 40 to 80 manipulation episodes in a productive hour, depending on task length and reset cost. To match even one percent of a web-scale text corpus in sample count, you need years of coordinated fleet time. That is exactly what the public record shows: DROID needed 13 institutions and 564 distinct scenes to gather 76,000 episodes (arXiv:2403.12945).
The practical implication: data strategy is now operations strategy. The teams shipping the strongest VLA and humanoid models are the ones running the most disciplined data factories, not the ones with the cleverest augmentation tricks.
Why Raw Hours Are Not Usable Hours
Usable hours are the QA-accepted portion of collected time, and the gap between raw and usable hours is the most underestimated number in robot data planning. Teams budget as if an 8-hour operator shift yields 8 hours of training data. It never does.
Here is where a shift actually goes, from our teleop cells:
| Shift component | Time lost (typical 8-hr shift) | Notes |
|---|---|---|
| Rig setup, calibration check | 30 to 45 min | Longer after any hardware change |
| Scene resets between episodes | 45 to 75 min cumulative | Task-dependent; deformable objects are worst |
| Operator breaks | 30 to 45 min | Non-negotiable; fatigue destroys quality |
| Rig downtime, retries | 15 to 40 min | Cable wear, gripper faults, dropped frames |
| QA rejection of recorded episodes | 10 to 25% of remainder | Occlusions, off-protocol grasps, sensor drift |
| Net usable output | 5.5 to 6.5 hours | 55 to 75% yield on wall-clock capture time |
Run the program-level math with the honest number. Ten rigs on two shifts at 6 usable hours each gives 120 usable hours per day, about 2,500 per month. Budget with 8 instead of 6 and you have overcommitted by a third before the first episode lands.
Yield is also not static. New operators start 10 to 15 points below tenured ones and converge over roughly six weeks. Every task or scene change dents yield temporarily. A program that changes its curriculum weekly and wonders why output swings is measuring the cost of its own churn.
The Bottleneck Moves as You Scale
Bottleneck migration is the pattern where each solved constraint exposes the next one downstream, and in robot data programs the sequence is predictable: rigs, then operators, then QA, then infrastructure.
- Rigs (weeks 0 to 8). Easiest to fix. An ALOHA-class bimanual cell costs $25k to $35k in our benchmarks and can be replicated in weeks.
- Operators (months 2 to 4). Hiring is easy; retention and protocol consistency are not. Plan 1.2 operators per rig-shift to absorb training and absence.
- QA (months 3 to 6). The silent killer. Reviewing one teleop hour takes 15 to 25 minutes with decent tooling. Without roughly 1 QA reviewer per 4 operators, a backlog forms, and defects like calibration drift go undetected for weeks, poisoning training runs retroactively.
- Infrastructure (months 4+). At 120 usable hours per day you ingest 2 to 4 TB daily. Ingest SLAs, LeRobot or RLDS formatting, dataset versioning, and storage tiering stop being nice-to-haves and become throughput constraints in their own right.
Teams that add rigs without staffing the downstream stages buy hardware that idles. We have audited programs where rig utilization sat below 50 percent because QA review, not capture, set the pace of the whole line.
What Actually Relieves the Bottleneck
Relieving the robot data bottleneck means increasing usable-hour throughput per dollar, and only three levers move that number materially.
- Raise yield before adding capacity. Moving yield from 60 to 72 percent on 10 rigs adds the output of 2 free rigs. Tenured operators, stable protocols, and same-day QA feedback are the mechanisms.
- Buy volume where it is cheaper than your floor. In-house fully loaded costs typically land at $42 to $64 per usable hour at Stage 3 scale; vendor rates run $28 to $60. Below roughly 1,500 usable hours per month, buying usually wins. The full breakeven model is in our complete guide to scaling robot data programs.
- Pool and standardize. Open X-Embodiment exists because 21 institutions shared a format. Standardizing on LeRobot or RLDS from day one means vendor data, partner data, and internal data merge instead of requiring conversion projects.
Simulation and human video help at the margins, especially for pretraining, but every strong manipulation result we have reproduced still ran through real teleoperated episodes at the fine-tuning stage. The bottleneck can be managed. It cannot yet be skipped.
Next Step
Next step: the full playbook, including the four-stage maturity model, hiring ratios, and the build-vs-buy breakeven sheet, is in The Complete Guide to Scaling Robot Data Programs. If you want to pressure-test your own throughput plan against our benchmarks, book a demo and bring your target usable-hours number.
Frequently Asked Questions
Why is data collection the main bottleneck in physical AI?
Because robot training data requires a physical event for every sample. Collection speed is capped by rigs, operators, and scene resets, so datasets grow linearly with operational investment while compute and model capacity grow much faster.
What percentage of collected robot data is actually usable?
In DexSet’s teleoperation programs, 55 to 75 percent of wall-clock capture time survives setup losses, resets, downtime, and QA rejection. Yield improves with operator tenure and stable protocols.
How much robot data do the largest public datasets contain?
Open X-Embodiment pools over 1 million trajectories from 21 institutions, DROID contains 76,000 episodes across 564 scenes from 13 institutions, and RT-1 used about 130,000 episodes gathered by 13 robots over 17 months.
Does simulation remove the robot data bottleneck?
Not yet. Simulation and egocentric human video reduce pressure at the pretraining stage, but current VLA and manipulation results still depend on real teleoperated episodes for fine-tuning, so physical collection capacity remains the binding constraint.
How do I calculate my robot data program’s real throughput?
Multiply rigs by shifts per day by usable hours per shift. A realistic usable-hours figure for an 8-hour teleop shift is 5.5 to 6.5, so 10 rigs on 2 shifts produce about 120 usable hours per day.
Enoch Pakanati
Enoch Pakanati is the strategic architect behind DexSet’s mission to become the undisputed market leader in robotics training data. He oversees the company’s growth strategy, focusing on capturing dominant market share across all data modalities required for modern robotics, including egocentric capture, teleoperation, and simulation-to-real data pipelines.
At DexSet, Enoch is responsible for transforming the company’s deep technical capabilities into a market-leading brand that foundation model labs and robotics OEMs trust implicitly. He focuses on scaling DexSet’s global footprint and ensuring the company stays ahead of the industry’s rapidly evolving data needs. His leadership is centered on one objective: making DexSet the singular, global standard for the data that powers the robotics revolution.