Comparing Approaches to Scaling Robot Data Programs: Pros, Cons, and Costs
For two weeks last spring, the blended cost per usable hour on a dashboard we share with one customer climbed from $44 to $57, and nobody on either side could explain it. Rig count had not changed. Wages had not changed. The eventual answer was source mix: the program’s in-house share had crept up while its vendor volume fell, and the two channels have genuinely different cost structures. Nobody had priced that difference explicitly, so nobody recognized it when it moved.
The same blind spot shows up in budget meetings, where three people quote three numbers for the same 15,000-hour data plan because each is quietly assuming a different collection approach. The spread is real, not a spreadsheet error. In-house collection, external vendors, pooled or public data, and hybrid models have genuinely different cost structures, ramp times, and failure modes. Vendors rarely publish rates, in-house programs rarely compute their fully loaded floor, and public datasets look free until you price the engineering to make them useful. So teams compare apples to invoices.
The thesis of this post: the four sourcing approaches differ less on sticker price than on ramp time, control, and failure mode, and a budget you can defend prices all four before committing to any. To make that concrete, this post puts all four approaches on one table with real numbers: cost per usable hour, ramp time, control, and where each one breaks. The in-house figures come from DexSet’s own operating costs; the vendor ranges are our published benchmarks; the public-data reference points are Open X-Embodiment, DROID, and RT-1.
By the end you should be able to defend a data budget to your board with one chart. For the full maturity model and hiring math behind these numbers, see The Complete Guide to Scaling Robot Data Programs.
Key Takeaways – Four approaches exist: in-house, vendor, public/pooled data, and hybrid. Nearly every program at scale ends up hybrid. – In-house fully loaded floors run $42 to $64 per usable hour at 10-rig scale; vendor teleop rates run $28 to $60. – In-house breakeven arrives near 1,500 sustained usable hours per month for 12+ months, and ramp takes 3 to 5 months. – Public datasets (Open X-Embodiment’s 1M+ trajectories, DROID’s 76k episodes) are strong for pretraining but rarely match your embodiment or task distribution. – Decide with three questions: volume, duration, and how much curriculum control you need.
What Are the Approaches to Scaling Robot Data?
Scaling approaches for robot data are the four sourcing models a team can use to grow usable-hour throughput: building an in-house collection operation, contracting external data vendors, adapting public or pooled datasets, and combining these in a hybrid program. Each trades cost, speed, control, and distribution match differently.
The choice is rarely permanent. Most teams we work with move through several: public data for pretraining experiments, a vendor for the first serious volume, a small in-house cell for curriculum iteration, then a hybrid steady state. The mistake is not picking a suboptimal approach; it is picking one without pricing the alternatives.
The Master Comparison Table
The table below compares the four approaches on the numbers that decide budgets. In-house costs assume a Stage 3 program (10 bimanual teleop rigs, 2 shifts, roughly 2,500 usable hours per month); all figures are DexSet benchmarks and planning ranges, not quotes.
| Factor | In-house | External vendor | Public / pooled data | Hybrid |
|---|---|---|---|---|
| Cost per usable hour | $42 to $64 fully loaded | $28 to $60 | $0 license; $15k to $60k one-time engineering to adapt | $35 to $55 blended |
| Ramp to full volume | 3 to 5 months | 2 to 6 weeks | Days to weeks | 4 to 8 weeks |
| Capex | $250k to $350k for 10 rigs | None | None | $50k to $120k |
| Curriculum control | Full | Contractual; days of turnaround | None | Full on internal cell |
| Embodiment match | Exact | Negotiable | Rarely exact | Exact where it matters |
| Scene diversity | Limited by your sites | High (vendor multi-site) | Very high (DROID: 564 scenes) | High |
| Scaling burst capacity | Poor (hiring-bound) | Good | N/A | Good |
| Main failure mode | QA under-staffing collapses yield | Distribution drift from your protocol | Domain gap to your robot | Calibration mismatch between sources |
In-House Collection: Full Control, Slow Ramp, Real Floor
In-house collection means owning rigs, operators, QA, and infrastructure yourself, and its true cost is the fully loaded floor, not the operator’s wage. Operator labor is only $12 to $16 of the $42 to $64 total; QA staffing (1 reviewer per 4 operators), rig amortization, data engineering, facilities, and storage make up the rest.
The pros are decisive when they matter. You get exact embodiment match, same-day curriculum changes, and data that never leaves your walls. Teams iterating fast on a proprietary humanoid have little alternative for their core loop.
The cons are ramp and fragility. Expect 3 to 5 months from budget approval to stable throughput, and note that RT-1 needed 13 robots and 17 months to reach roughly 130,000 episodes (arXiv:2212.06817). Output also collapses quietly when downstream stages lag; a QA backlog can idle half a rig fleet.
Choose in-house when volume exceeds roughly 1,500 usable hours per month, the need lasts 12+ months, and curriculum control is worth paying for.
External Vendors: Fast Volume, Contract Discipline Required
Vendor collection means paying a specialist per usable hour, typically $28 to $60 for teleoperation in our benchmarks, with egocentric capture materially cheaper. The vendor carries recruiting, yield management, QA calibration, and rig maintenance; you carry the specification.
Speed is the headline advantage: weeks to volume instead of months, no capex, and burst capacity when a training milestone moves up. Scene diversity is often better than any single in-house site can offer.
The risks concentrate in specification quality. If your acceptance criteria are vague, you will receive episodes that pass the vendor’s QA and fail your training run. Insist on a shared golden set (we calibrate on about 200 episodes), per-episode metadata in LeRobot or RLDS format, and yield reporting in usable hours, not raw hours. A vendor quoting raw hours is quoting a number 25 to 45 percent larger than what you will train on.
Choose a vendor when you need volume before you could possibly hire for it, or when your total need sits below the in-house breakeven.
Public and Pooled Datasets: Free Data, Paid Integration
Public datasets are shared corpora like Open X-Embodiment (1M+ trajectories across 21 institutions, arXiv:2310.08864) and DROID (76k episodes, 564 scenes, 13 institutions, arXiv:2403.12945), and their license cost of zero hides a real engineering cost to adapt them. Budget weeks of data engineering for format normalization, camera-frame conventions, action-space mapping, and filtering, which is why we put $15k to $60k of one-time cost on the table above.
They are excellent for pretraining and for benchmarking your own data quality. They are poor as a sole source: the embodiment, camera placement, and task distribution almost never match yours, and no amount of filtering manufactures episodes of your robot doing your tasks.
Use public data as a floor under your pretraining, and as a reason to standardize your own formats early so everything merges cleanly.
Hybrid: What Programs at Scale Actually Run
A hybrid program keeps a small in-house cell for iteration while vendors carry volume, and it is where nearly every Stage 4 operation lands. The internal cell (2 to 4 rigs) owns curriculum experiments and embodiment-critical tasks. The vendor side owns bulk hours and scene diversity. Public data anchors pretraining.
The blended cost typically lands at $35 to $55 per usable hour, and the failure mode is calibration mismatch: two sources whose “acceptable episode” differs. The fix is procedural, not technical: one shared golden set, one format, one metadata schema, joint yield reviews. Open X-Embodiment demonstrated at field scale that pooling works when the format discipline holds.
Decision Checklist
- Under 500 total usable hours needed: buy, or adapt public data. Never build.
- Uncertain duration or volume: vendor, with a small internal cell only if embodiment iteration is core.
- 1,500+ usable hours per month for 12+ months: build toward in-house, keep vendor burst capacity.
- Any multi-source setup: shared golden set, LeRobot or RLDS format, usable-hour (not raw-hour) contracts.
Next Step
Go deeper: the maturity model, throughput calculator, and full breakeven worksheet behind this comparison live in The Complete Guide to Scaling Robot Data Programs. Or book a demo and we will run the breakeven math against your actual volume target.
Frequently Asked Questions
How much does robot data collection cost per hour?
Vendor teleoperation data runs $28 to $60 per usable hour in DexSet’s benchmarks, depending on rig type, task complexity, and QA depth. Fully loaded in-house costs at 10-rig scale run $42 to $64 per usable hour.
When does building in-house robot data collection beat buying?
At sustained volumes above roughly 1,500 usable hours per month lasting 12 or more months. Below that, vendor rates undercut the in-house floor and avoid a 3-to-5-month ramp and $250k+ rig capex.
Are public robot datasets good enough to train a VLA model?
They are strong for pretraining: Open X-Embodiment offers over 1 million trajectories and DROID adds 76,000 episodes across 564 scenes. But embodiment and task mismatch means fine-tuning still requires data collected on or for your own platform.
What should a robot data vendor contract specify?
Usable hours rather than raw hours, a shared golden set for QA calibration, per-episode metadata, delivery in LeRobot or RLDS format, and yield reporting. Raw-hour contracts overstate deliverable data by 25 to 45 percent.
What is a hybrid robot data program?
A model where a small in-house cell (2 to 4 rigs) handles curriculum iteration and embodiment-critical tasks while vendors supply volume and scene diversity, blending to roughly $35 to $55 per usable hour.
Sainath Gupta
Sainath Gupta is the visionary leader of DexSet, a company dedicated to building the foundational data layer for the robotics revolution. Sainath recognized early on that the primary bottleneck for physical AI wasn't hardware or model architecture, but the lack of high-quality, scalable training data.
At DexSet, he has pioneered a "manufacturing-first" approach to data collection. This includes the development of transparent cost-per-hour benchmarks and the scaling of global networks for egocentric and teleoperation capture. Sainath is a vocal advocate for the industry’s shift toward treating human experience as the primary pretraining substrate for robots. His leadership at DexSet is focused on one goal: providing the millions of hours of high-fidelity data required to bring humanoid robots out of the lab and into the real world.