The Complete Guide to Scaling Robot Data Programs (2026)
Two manipulation teams we audited last year bought nearly identical hardware in the same quarter: eight bimanual teleop rigs each. Eighteen months on, one was delivering around 1,800 usable hours a month at better than 90 percent QA acceptance. The other was producing under 400 and re-collecting a third of what it shipped. The rigs were not the difference. The first team had built shifts, a QA organization, and an ingest SLA before it added a fifth rig; the second had cloned hardware and hoped.
That contrast is the thesis of this guide: robot data scale is an operations achievement, not a hardware purchase. Robot data does not scale like text or images. You cannot scrape it. Every episode requires a physical rig, a human operator, a real scene, and a QA pass. The teams that trained the reference models did not solve this with clever algorithms; they solved it with operations. Google’s RT-1 dataset took around 130,000 episodes, 13 robots, and 17 months of continuous collection (arXiv:2212.06817). DROID took 13 institutions coordinating 76,000 episodes across 564 scenes (arXiv:2403.12945). Open X-Embodiment pooled over one million trajectories, and it still needed 21 institutions to get there (arXiv:2310.08864).
This guide gives you the operational playbook: a four-stage maturity model for robot data programs, the throughput math that predicts your real output, a build-versus-buy breakeven framework, hiring ratios, and the infrastructure decisions that keep a petabyte-scale program sane. We run these programs daily at DexSet, and the numbers below come from our own rigs, pipelines, and cost sheets.
TL;DR – A robot data program is the combination of rigs, operators, QA, and infrastructure that produces training-ready episodes on a predictable schedule. – Programs mature through four stages: ad-hoc lab capture, dedicated rigs, production data factory, and multi-site plus external vendors. – Output is governed by simple math: rigs × shifts × usable hours per shift. 10 rigs × 2 shifts × 6 usable hours = 120 usable hours per day. – Raw hours are not usable hours. Expect 55 to 75 percent yield after QA; plan capacity against usable hours only. – In-house cost floors typically land at $40 to $65 per usable hour fully loaded. Vendors range from $28 to $60. Breakeven usually sits near 1,500 sustained usable hours per month for 12 or more months. – Staff with ratios: roughly 1 QA reviewer per 4 operators, 1 data engineer per 8 to 10 rigs, 1 site lead per shift.
What Is a Robot Data Program?
A robot data program is an ongoing operational system that produces robot learning data (teleoperation episodes, egocentric video, exocentric multi-view capture) at a defined quality bar and a defined rate. It is a program, not a project: it has staffing, throughput targets, QA acceptance criteria, and a budget line, the same way a manufacturing line does.
The distinction matters because most teams treat data collection as a task (“collect 500 pick-and-place demos”) rather than a capability (“produce 100 usable hours per week at over 95 percent QA acceptance”). Tasks end. Programs compound. VLA models trained on Open X-Embodiment scale benefited from data collected continuously over years, not from one heroic sprint.
Scaling a robot data program means growing that capability along four axes at once: capture capacity (rigs and operators), quality assurance (review throughput and acceptance criteria), infrastructure (ingest, storage, versioning), and coordination (scheduling, task curricula, scene rotation). Neglect any one axis and the others stall. We have watched teams double their rig count and gain almost nothing because QA became the choke point within two weeks.
The Robot Data Program Maturity Model
The Robot Data Program Maturity Model is a four-stage framework that classifies data operations by structure, throughput, and staffing, from ad-hoc lab capture to multi-site production. We built it after auditing dozens of customer programs, because “how much data can you make?” turns out to be answerable almost entirely from which stage you are in.
| Dimension | Stage 1: Ad-hoc lab capture | Stage 2: Dedicated rigs | Stage 3: Production data factory | Stage 4: Multi-site + vendors |
|---|---|---|---|---|
| Who collects | Researchers, between other work | Part-time operators on fixed rigs | Full-time operators in shifts | Internal shifts plus external vendors |
| Rigs | 1 to 2, shared | 2 to 5, dedicated | 8 to 20, standardized | 20+, multiple locations |
| Typical output | 5 to 20 usable hr/month | 60 to 150 usable hr/month | 1,500 to 3,000 usable hr/month | 4,000+ usable hr/month |
| QA | Collector self-checks | Spot checks, no dedicated staff | Dedicated QA org, acceptance SLAs | Cross-site QA calibration |
| Cost visibility | None | Rough | Cost per usable hour tracked weekly | Blended internal vs vendor $/hr |
| Failure mode | Data drought after paper deadline | Operator churn, drifting protocols | Yield collapse if QA under-staffed | Distribution mismatch across sites |
Stage 1: Ad-hoc lab capture. Researchers collect their own demonstrations when they need them. Fine for prototyping ACT or ALOHA-style policies (arXiv:2304.13705), fatal for foundation models, because output stops whenever the researchers get busy.
Stage 2: Dedicated rigs with part-time operators. The first real investment: fixed rigs, written task protocols, a few trained operators. Most seed-stage robotics startups live here. The trap is treating it as scalable; part-time operators churn, and protocols drift without a QA function.
Stage 3: Production data factory. Shifts, throughput SLAs, a QA org with written acceptance criteria, and a data engineering function owning ingest and versioning. This is where cost per usable hour becomes a managed metric. RT-1’s collection operation, 13 robots running for 17 months, is the canonical public example of Stage 3 discipline.
Stage 4: Multi-site plus external vendors. Once internal capacity saturates, you add sites or vendors. The hard problem shifts from throughput to distribution control: DROID deliberately spread collection across 564 scenes and 13 institutions to get visual diversity, and cross-site programs need the same intentionality, plus calibration so that two sites’ “acceptable episode” means the same thing.
Locate yourself honestly. The most expensive mistake we see is a Stage 2 team signing a Stage 3 data commitment to their investors.
Core Modalities and What Scale Demands From Each
Modality choice determines your rig cost, your storage bill, and your yield, so a scaling plan has to be modality-specific. The main options:
- Teleoperation (leader-follower or VR): The workhorse for manipulation. Highest cost per hour, highest action-label quality. Entity chain: teleoperation → ALOHA-style rigs → imitation learning → VLA models.
- Egocentric human video (head-mounted, mono or stereo): Cheapest to scale, no robot embodiment, increasingly used for pretraining following Ego4D and EgoExo4D (arXiv:2311.18259).
- Exocentric multi-view: Fixed external cameras around a workspace. Moderate cost, essential for scene-level context and world-model work.
- Stereo and depth: Adds calibration burden per rig; at 20 rigs, calibration drift becomes a standing QA line item, not an occasional annoyance.
| Modality | Rig cost (typical) | Raw data rate | Usable-hour yield | Best for |
|---|---|---|---|---|
| Bimanual teleop (ALOHA-class) | $25k to $35k | 15 to 30 GB/hr (3 cams, 480p to 720p, 30 fps) | 55 to 70% | Manipulation policies, VLA fine-tuning |
| VR teleop (humanoid) | $8k to $20k + robot | 20 to 40 GB/hr | 50 to 65% | Whole-body, humanoid |
| Egocentric stereo | $1k to $5k per wearer | 8 to 15 GB/hr | 70 to 85% | Pretraining, hand-object interaction |
| Exocentric 4-cam | $4k to $10k | 25 to 50 GB/hr | 75 to 90% | Scene context, world models |
Rig costs and yields above are DexSet benchmarks from our own fleet; treat them as planning ranges, not quotes. For an external anchor, the ALOHA reference build is about $20k in hardware alone (arXiv:2304.13705); our cell figures include cameras, compute, and fixturing on top of the arms.
Throughput Math: Rigs × Shifts × Usable Hours
Throughput math is the capacity model that converts rigs, shifts, and yield into a predictable output number, and it is the single most useful planning tool in this guide. The formula:
Usable hours per day = rigs × shifts per day × usable hours per operator-shift
The third term does the damage. An 8-hour shift never produces 8 usable hours. Subtract setup and calibration (30 to 45 minutes), scene resets between episodes, operator breaks, rig downtime, and then the QA rejection rate on top. In our teleop operations, a well-run 8-hour shift nets 5.5 to 6.5 usable hours. Call it 6.
So a Stage 3 program with 10 rigs running 2 shifts:
10 rigs × 2 shifts × 6 usable hours = 120 usable hours per day, roughly 2,500 usable hours per month at 21 working days.
Now sanity-check any collection commitment against public reference points. RT-1’s ~130,000 episodes over 17 months with 13 robots averages out to under 600 episodes per robot per month, and Google ran a disciplined operation. If your plan implies triple that per rig with no QA org, the plan is wrong, not ambitious.
Two corollaries we hold operators to:
- Plan against usable hours, never raw hours. A 70 percent yield means a 100-hour commitment requires 143 raw hours of capacity.
- Yield is a managed variable. Yield rises with operator tenure (we see 10 to 15 point improvements over an operator’s first six weeks) and falls every time you change tasks, scenes, or acceptance criteria. Budget a yield dip for every curriculum change.
Cost, Economics, and the Build-vs-Buy Breakeven
Build-versus-buy analysis for robot data compares your fully loaded in-house cost per usable hour against vendor pricing at your required volume and duration. Both numbers are knowable; most teams compute neither.
The in-house cost floor. Add up everything per usable hour at Stage 3 scale (10 rigs, 2 shifts, ~2,500 usable hr/month):
| Cost component | Basis | $/usable hour |
|---|---|---|
| Operator labor | ~$28/hr loaded, at 75% yield | $12 to $16 |
| QA reviewers | 1 per 4 operators, ~$32/hr loaded | $6 to $9 |
| Rig amortization | $30k rig over 24 months, ~250 usable hr/mo | $4 to $6 |
| Data engineering | 2 engineers per 10 rigs, ~$16k/mo each loaded | $10 to $14 |
| Facility, power, consumables, props | warehouse space, scene materials | $4 to $7 |
| Storage and compute (ingest, transcode) | see infrastructure section | $2 to $4 |
| Management overhead | site lead per shift, program mgmt | $4 to $8 |
| Total in-house floor | $42 to $64 |
Vendor pricing for comparable teleop data, in our benchmarks, runs $28 to $60 per usable hour depending on rig type, task complexity, and QA depth. Egocentric capture is cheaper on both sides.
The breakeven logic:
- Under ~500 usable hours total: Buy. Your rig capex and hiring lead time never amortize.
- 500 to 1,500 usable hours per month, or uncertain duration: Buy, or run a small internal Stage 2 cell for rapid iteration while a vendor carries volume.
- 1,500+ usable hours per month sustained for 12+ months: Building approaches or beats vendor pricing, and you gain direct control of the task curriculum. Ramp time is the hidden cost: expect 3 to 5 months from budget approval to full Stage 3 throughput.
- Stage 4 reality: Almost every program at scale ends up hybrid. Internal rigs handle proprietary embodiments and fast iteration; vendors handle volume, scene diversity, and burst capacity.
One honest caveat, and this is opinion from the vendor side of the table: the strongest argument for building is not cost, it is iteration speed on your own embodiment. The strongest argument for buying is that yield, QA calibration, and operator management are someone else’s learned-the-hard-way problem.
The Hiring Plan: Ratios That Hold Up
A robot data hiring plan defines the roles and staffing ratios required to keep capture, QA, and infrastructure in balance as rig count grows. The ratios we run and recommend:
- Operators: 1.2 per rig-shift (the 0.2 covers breaks, training, and absence). 10 rigs on 2 shifts means about 24 operators.
- QA reviewers: 1 per 4 operators. QA review of a teleop hour takes 15 to 25 minutes with good tooling; under-staff this and unreviewed episodes pile up until a bad-calibration week silently poisons a training run.
- Data engineers: 1 per 8 to 10 rigs, minimum 2 once you are Stage 3. They own ingest, format conversion, versioning, and the QA tooling itself.
- Site lead: 1 per shift. Owns the schedule, scene rotation, and rig maintenance triage.
- Program manager: 1 per site once you pass roughly 15 operators.
Sequence matters as much as ratios. Hire the first QA reviewer before the fifth operator, and the first data engineer before the tenth. Every program we have seen invert that order paid for it in re-collected data.
Infrastructure: Ingest, Formats, Versioning, Storage
Data infrastructure for a scaled robot program is the pipeline that moves episodes from rig to training-ready dataset with provenance intact. At 120 usable hours per day of multi-camera teleop, you are ingesting 2 to 4 TB daily, and infrastructure choices become budget lines.
Ingest. Rigs write locally, then upload on a schedule with checksums and automatic retry. Same-day ingest is a real SLA: QA cannot review what has not landed, and yield problems you discover a week late cost a week of bad data.
Formats. Standardize early on LeRobot dataset format (github.com/huggingface/lerobot) or RLDS (github.com/google-research/rlds). Open X-Embodiment’s pooling of 21 institutions’ data was possible because RLDS gave everyone a common episode structure. Your future self, merging vendor data with internal data at Stage 4, will thank you.
Versioning. Treat datasets like code releases: immutable snapshots, semantic versions, and a manifest recording rig ID, operator ID, calibration state, task, scene, and QA verdict for every episode. When a training run regresses, per-episode provenance is the difference between a one-day diagnosis and a one-month one.
Storage economics. At S3-class pricing near $23 per TB-month for hot storage (aws.amazon.com/s3/pricing), a program producing 3 TB/day accumulates ~90 TB/month: roughly $2,000/month hot, compounding as long as you keep everything hot. Tier aggressively: raw footage to cold storage (about $1 to $4 per TB-month) after transcode, training-ready compressed episodes hot. Compression from raw to training format typically cuts volume 3 to 5×.
What the Public Mega-Datasets Teach About Scale
The public record on large robot datasets is short but consistent: scale came from organization, not from any single site getting faster.
- RT-1 (arXiv:2212.06817): ~130k episodes, 13 robots, 17 months. Lesson: fleet consistency and duration beat burst collection.
- DROID (arXiv:2403.12945): 76k episodes, 564 scenes, 13 institutions on a standardized Franka rig. Lesson: standardize the rig, diversify the scenes.
- Open X-Embodiment (arXiv:2310.08864): 1M+ trajectories pooled across 21 institutions. Lesson: common formats and metadata are what make pooled scale possible at all.
Every one of these is a Stage 3 or Stage 4 operation. None of them is a lab that simply worked harder.
Case Study Proof: A Humanoid Foundation Model Team’s Ramp
A humanoid foundation model team came to us at Stage 2: four VR teleop rigs, part-time operators, about 90 usable hours per month, and a training roadmap that needed 1,200 per month within two quarters. We ran their volume on DexSet capture cells while their internal team kept two rigs for curriculum iteration. Combined output crossed 1,300 usable hours per month in 11 weeks, with QA acceptance holding above 93 percent after we calibrated both sides on a shared 200-episode golden set. Their blended cost landed at $41 per usable hour, below their projected internal floor of $55 at that volume. The number that mattered most in hindsight was not throughput; it was the two-week-earlier detection of a wrist-camera calibration drift, caught because same-day QA was in the SLA.
Downloadable: Robot Data Capacity Planning Worksheet
We packaged the math in this guide, the maturity model self-assessment, throughput calculator, build-vs-buy breakeven sheet, and hiring ratio table, into a capacity planning worksheet plus a vendor RFP scorecard. If you are scoping next year’s data budget, start there.
Where to Go Deeper
Next Step
Ready to plan your ramp? Download the Robot Data Capacity Planning Worksheet and RFP scorecard, or book a demo to see DexSet capture cells, QA pipelines, and sample teleop and egocentric datasets. If you bring your target usable-hours number, we will walk through the breakeven math against your own cost assumptions.
Frequently Asked Questions
What is a robot data program?
A robot data program is an ongoing operation of rigs, operators, QA reviewers, and data infrastructure that produces robot learning data at a defined quality bar and rate, measured in usable hours per month rather than one-off collection tasks.
How many usable hours can one teleoperation rig produce per month?
A dedicated rig running two 8-hour shifts typically nets 220 to 280 usable hours per month, since each shift yields 5.5 to 6.5 usable hours after setup, resets, downtime, and QA rejections.
When should a robotics team build in-house data collection instead of buying?
Building starts to beat buying at roughly 1,500 sustained usable hours per month for 12 or more months, where in-house floors of $42 to $64 per usable hour compare against vendor rates of $28 to $60. Below that volume, or with uncertain duration, buying or a hybrid model is usually cheaper.
What staffing ratios does a scaled robot data program need?
Plan on about 1.2 operators per rig-shift, 1 QA reviewer per 4 operators, 1 data engineer per 8 to 10 rigs, and 1 site lead per shift.
How much storage does large-scale robot data collection require?
Multi-camera teleoperation generates 15 to 30 GB per raw hour, so a 120-usable-hour-per-day program ingests 2 to 4 TB daily. At about $23 per TB-month for hot cloud storage, tiering raw footage to cold storage after transcode is essential.
How large were RT-1, DROID, and Open X-Embodiment?
RT-1 used about 130,000 episodes collected over 17 months with 13 robots. DROID contains 76,000 episodes across 564 scenes from 13 institutions. Open X-Embodiment pooled more than 1 million trajectories across 21 institutions.
Sainath Gupta
Sainath Gupta is the visionary leader of DexSet, a company dedicated to building the foundational data layer for the robotics revolution. Sainath recognized early on that the primary bottleneck for physical AI wasn't hardware or model architecture, but the lack of high-quality, scalable training data.
At DexSet, he has pioneered a "manufacturing-first" approach to data collection. This includes the development of transparent cost-per-hour benchmarks and the scaling of global networks for egocentric and teleoperation capture. Sainath is a vocal advocate for the industry’s shift toward treating human experience as the primary pretraining substrate for robots. His leadership at DexSet is focused on one goal: providing the millions of hours of high-fidelity data required to bring humanoid robots out of the lab and into the real world.