Case Study: How We Scaled a Robot Data Program for a VLA Model
The budget meeting that started this program stalled on a single line item. The client’s CFO wanted to know why our proposal priced roughly 17,000 raw capture hours to deliver 12,000 usable ones, and whether the 5,000-hour gap was padding. It was yield, not padding, and the fact that nobody on their side had ever asked the question internally told us more about their situation than the rig audit did. Their model roadmap called for a fine-tuning corpus of roughly 12,000 usable teleoperation hours across 40 manipulation task families. Their internal operation, four bimanual rigs staffed by part-time operators, was producing 90 usable hours a month. At that rate the corpus would arrive in 2037.
The problem was not effort. Their operators worked hard and their rigs were well built. The problem was that they were running a Stage 2 operation against a Stage 3 requirement: no shifts, no dedicated QA, no ingest SLA, and no idea what their cost per usable hour actually was. When we computed it with them, it came out near $85, mostly because rig utilization sat at 31 percent and one in four episodes failed their own training team’s spot checks.
This post walks through the 11-week ramp that took the combined program to 1,300 usable hours per month at a blended $41 per usable hour: what we changed week by week, the yield curve, the QA calibration process, and the two mistakes that cost us time. Numbers are rounded but real, taken from the weekly ops reviews we still run with this team. The thesis those numbers argue: capacity additions did not rescue this program, shorter feedback loops did. The ramp worked because we split iteration from volume, calibrated QA on a shared golden set, and made review same-day.
If you want the general framework this ramp followed, it is in The Complete Guide to Scaling Robot Data Programs. This is what the framework looks like when it hits a deadline.
Key Takeaways – Baseline: 4 rigs, 90 usable hr/month, ~$85 per usable hour, 31% rig utilization. – Outcome at week 11: 1,300 usable hr/month combined, $41 blended cost per usable hour, QA acceptance above 93%. – The split that worked: DexSet cells carried volume (12 rigs, 2 shifts) while the client kept 2 rigs for curriculum iteration. – A 200-episode shared golden set aligned QA between both sides in week 2 and prevented the classic multi-source drift problem. – Biggest single win was boring: same-day QA feedback lifted operator yield from 58% to 74% in five weeks.
The Starting Point: A Stage 2 Program With Stage 3 Commitments
A Stage 2 robot data program has dedicated rigs and trained part-time operators but no shift structure, QA organization, or throughput SLA, and that was exactly the client’s position. Their four ALOHA-class rigs were capable hardware. The operation around them was the constraint.
The baseline audit told the story in four numbers:
| Metric | Baseline (week 0) |
|---|---|
| Usable hours per month | 90 |
| Rig utilization (capture time / available time) | 31% |
| QA acceptance rate | 76%, measured by ad-hoc spot checks |
| Effective cost per usable hour | ~$85 |
Utilization was low because operators were part-time and unscheduled; rigs sat idle most of every day. Acceptance suffered from protocol drift: three operators had three different ideas about what a completed fold, pour, or insertion looked like, because nothing was written down. And the training team was consuming episodes weeks after capture, so defects were discovered long after hundreds of similar episodes had been recorded.
None of this is unusual. It is the default state of a lab-grown collection effort asked to feed a foundation model. The relevant reference point is that Google’s RT-1 corpus, about 130,000 episodes, took 13 robots and 17 months of organized fleet collection (arXiv:2212.06817). Corpora of that class do not come from unscheduled rigs.
The Ramp Plan: Split Iteration From Volume
The core design decision was a hybrid split: the client’s internal cell would own curriculum iteration while DexSet cells owned volume. The client kept 2 of their 4 rigs for fast experiments on new task families. We stood up 12 bimanual cells on our floor, matched to their camera and action-space conventions, running two 8-hour shifts.
The throughput math, which we committed to as an SLA, worked out as follows: 12 rigs × 2 shifts × 6 usable hours gives 144 usable hours per day at steady state. We ramped to that over eight weeks rather than promising it on day one, because yield always starts low with new operators and a new curriculum.
Week by week, the ramp looked like this:
| Weeks | Focus | Usable hr/month (run rate) |
|---|---|---|
| 1 to 2 | Rig matching, golden set QA calibration, 8 operators trained | 180 |
| 3 to 4 | First shift at full strength, ingest SLA live, LeRobot pipeline validated | 420 |
| 5 to 7 | Second shift hired and trained, QA org at 1 per 4 operators | 780 |
| 8 to 9 | Yield push: same-day QA feedback loop, protocol revisions frozen | 1,050 |
| 10 to 11 | Steady state plus client’s 2 iteration rigs | 1,300 |
Staffing followed the ratios we use everywhere: 29 operators across both shifts (about 1.2 per rig-shift), 7 QA reviewers, 2 data engineers, and a site lead per shift.
QA Calibration: The 200-Episode Golden Set
A golden set is a fixed collection of reference episodes that both data producer and consumer score independently to align acceptance criteria, and building one in week 2 was the highest-return hour-for-hour work of the whole program. We captured 200 episodes spanning all 40 task families, including deliberately marginal cases: partial insertions, occluded grasps, near-boundary camera framing.
Both QA teams scored all 200 blind. First-pass agreement was 81 percent, which sounds acceptable and is not; at 1,300 hours a month, 19 percent disagreement is 250 disputed hours. The disagreements clustered in three areas (retry handling, reset completeness, and what counted as an occlusion), so the written protocol got three new pages and we rescored. Agreement hit 96 percent, and it held through the ramp because we rechecked a 20-episode sample every two weeks.
Every episode shipped with per-episode metadata: rig ID, operator ID, calibration state, task family, scene ID, QA verdict, and protocol version, packaged in LeRobot format (github.com/huggingface/lerobot). When the client later merged our data with their internal cell’s output and a slice of Open X-Embodiment for pretraining (arXiv:2310.08864), the merge was a config change, not a project.
The Yield Curve: 58 to 74 Percent in Five Weeks
Yield is the fraction of wall-clock capture time that survives QA, and it is the number that decided this program’s economics. Our new operators started at 58 percent in weeks 3 and 4. By week 9 the fleet averaged 74 percent. Nothing exotic drove the improvement; three boring mechanisms did.
First, same-day QA feedback. Operators saw their rejected episodes, with reasons, before their next shift. Rejection causes fell fastest in the categories operators could directly control: off-protocol grasps dropped by two-thirds.
Second, protocol freezes. We batched all task-definition changes into a Monday release instead of letting them trickle in daily. Every curriculum change still dented yield for a few days, so we scheduled changes against the SLA rather than pretending they were free.
Third, calibration as a shift ritual. A five-minute fiducial check at shift start caught camera drift twice in the ramp, both times before a full shift of episodes was affected. The client’s earlier operation had once lost three weeks of data to exactly this failure, discovered only at training time.
Results and the Two Mistakes Worth Copying From
At week 11 the combined program hit its numbers: 1,300 usable hours per month, 93 to 95 percent QA acceptance against the golden set, blended cost of $41 per usable hour, and the client’s fine-tuning corpus on schedule for delivery in nine months instead of eleven years. Their VLA’s task-family success rates on held-out evaluations improved as data volume crossed each curriculum milestone, though model results are theirs to publish, not ours.
Two mistakes deserve daylight. We initially under-staffed QA at 1 reviewer per 6 operators to save cost; by week 5 the review backlog hit four days and we were flying blind on yield, so we hired to 1 per 4 and the backlog cleared in a week. And we let one high-variance task family (deformable bag packing) into the week 3 curriculum before protocols were stable; its 40 percent rejection rate dragged fleet yield down and demoralized new operators. We pulled it, stabilized the protocol on the client’s iteration rigs, and reintroduced it in week 8 at 71 percent yield. Sequence hard tasks late. The DROID team’s choice to standardize one rig platform across 13 institutions and 564 scenes reflects the same instinct: control variance where you can, spend it where it buys diversity (arXiv:2403.12945).
Next Step
See it in person: the capture cells, golden-set QA process, and LeRobot delivery pipeline from this case study are all demoable. Book a demo, or start with the full playbook in The Complete Guide to Scaling Robot Data Programs.
Frequently Asked Questions
How long does it take to scale a robot data program?
This program went from 90 to 1,300 usable hours per month in 11 weeks using a hybrid model. A fully in-house ramp typically takes 3 to 5 months from budget approval to stable Stage 3 throughput.
How much did the data cost per usable hour?
The blended cost landed at $41 per usable hour, against a baseline of roughly $85 in the client’s original low-utilization operation and a projected in-house floor of $55 at their target volume.
What is a golden set in robot data QA?
A fixed set of reference episodes, 200 in this program, that producer and consumer QA teams score independently to align acceptance criteria. Agreement improved from 81 to 96 percent after one protocol revision cycle.
What yield should a teleoperation program expect?
New operators in this program started near 58 percent usable yield and reached a 74 percent fleet average by week 9, driven by same-day QA feedback, batched protocol changes, and shift-start calibration checks.
How much teleoperation data does a VLA model need?
It varies with scope. This client targeted 12,000 usable hours across 40 task families for fine-tuning on top of public pretraining data. For reference, RT-1 used about 130,000 episodes and Open X-Embodiment pools over 1 million trajectories.
Enoch Pakanati
Enoch Pakanati is the strategic architect behind DexSet’s mission to become the undisputed market leader in robotics training data. He oversees the company’s growth strategy, focusing on capturing dominant market share across all data modalities required for modern robotics, including egocentric capture, teleoperation, and simulation-to-real data pipelines.
At DexSet, Enoch is responsible for transforming the company’s deep technical capabilities into a market-leading brand that foundation model labs and robotics OEMs trust implicitly. He focuses on scaling DexSet’s global footprint and ensuring the company stays ahead of the industry’s rapidly evolving data needs. His leadership is centered on one objective: making DexSet the singular, global standard for the data that powers the robotics revolution.