Case Study: What a Year-Long VLA Data Program Taught Us About 2027
We enforce an internal rule at DexSet that surprises visiting customers: no data program keeps its kickoff modality mix past the first quarterly ablation review without re-justifying it from evidence. This retrospective is about the program that taught us the rule. A VLA training program does not fail on day one. It fails in month four, when the first big fine-tuning run comes back flat and nobody can say whether the problem is the model, the mix, or the data. We spent twelve months inside exactly that situation with one customer, a robot foundation model team we will keep anonymous, and the program ended somewhere very different from where it started: different modality mix, different pricing structure, different QA gates, and a 45 percent drop in cost per successful demonstration.
Programs drift like this because the ground shifts under them. When this one kicked off in mid-2025, the customer’s plan reflected the best public evidence of the time: teleoperation-heavy, priced per raw hour, QA as a spot-check. Then their own ablations, the same forces we document in The State of Robotics Training Data 2027, pushed them toward human video at scale, usable-hour contracts, and audit-grade provenance. A year of purchase orders is a truth serum for data strategy, and the thesis this year left us with is simple: contract structure and evidence-driven mix reviews, not collection volume, determine whether a data program compounds or leaks.
This retrospective shares what we can: the mix shift quarter by quarter, the rejection-rate story that changed the commercial terms, the compliance audit that arrived mid-program, and the numbers we would plan around for 2027. Per our anonymization agreement, we name no company, no model, and no revenue figures. Everything else is first-hand, Level 1 evidence from our side of the program.
If you are scoping a 2027 data program, treat this as one team’s trace through the decisions you are about to make.
Key Takeaways – Over 12 months the program delivered roughly 2,100 usable teleop hours and 38,000 usable egocentric video hours; the spend ratio flipped from 80/20 teleop-heavy to 35/65 video-heavy. – Early teleop batches hit a 22 percent QA rejection rate under raw-hour terms; moving to usable-hour contracts with a written rubric cut disputes to near zero and rejection to roughly 10 percent within two months. – A mid-program legal audit demanded consent provenance for every identifiable person in the egocentric corpus. Episode-keyed consent records turned a potential months-long crisis into a five-day export. – Cost per successful task demonstration fell about 45 percent from Q1 to Q4, driven by mix shift and the QA feedback loop, not by cutting operator rates. – The customer’s internal 40-task evaluation success rate roughly doubled; we credit data for the cost curve and stay agnostic on how much of the model gain was data versus architecture.
The Program as Designed: Mid-2025 Assumptions
The original program design was a teleoperation-first plan: 80 percent of budget on bimanual and mobile-manipulation teleop, 20 percent on egocentric video as a hedge, priced per raw hour with acceptance handled informally. That design was defensible when it was written. ALOHA-line results had made teleop the proven currency of imitation learning (arXiv:2304.13705), and the customer’s early prototypes were trained almost entirely on embodiment-matched demonstrations.
Quarter one ran to plan mechanically: rigs shipped, operators ramped, hours flowed. Two problems surfaced anyway. The first was quality, covered in the next section. The second was strategic: the customer’s pretraining team was running ablations on public mixtures, including Open X-Embodiment (arXiv:2310.08864) and human video corpora in the Ego4D and EgoExo4D style (arXiv:2110.07058; arXiv:2311.18259), and the ablations kept saying the same thing: video scale moved downstream success more per dollar than marginal teleop hours did, once a teleop floor was met.
The Rejection-Rate Story: How 22 Percent Changed the Contract
A QA rejection rate is the fraction of delivered hours that fail acceptance review, and in this program it started at 22 percent for teleop. That number needs honest context: some of it was on us (two stations with intermittent camera desync), some was task-definition ambiguity (what counts as “done” for shelf restocking was underspecified), and some was operator behavior that only became visible under review, like unlabeled error recoveries.
Under raw-hour terms, every rejected batch became a negotiation. By month three both sides wanted out of that dynamic, and the fix became the template we now use everywhere: usable-hour pricing against a written acceptance rubric. The rubric specified task success per task family, a 10 ms stream-sync tolerance, calibration reprojection thresholds, and a label-accuracy sampling plan. Crucially, rejection reasons flowed back to my rig operators weekly. Rejection fell to roughly 10 percent within two months and stabilized between 10 and 12 percent for the rest of the program. The lesson we carry into 2027: rejection rates are not weather. They respond to feedback loops, and the contract structure determines whether the loop exists.
The Mix Shift, Quarter by Quarter
The defining trend of the program was budget migrating from teleoperation to egocentric human video as the customer’s own evidence accumulated. The table below shows the shift in new-order spend and the unit economics behind it, from our delivery records.
| Quarter | Teleop share of new spend | Egocentric share | Blended cost per usable hour (teleop) | Blended cost per usable hour (egocentric) | Cost per successful demo (indexed, Q1 = 100) |
|---|---|---|---|---|---|
| Q1 (2025) | 80% | 20% | $59 | $30 | 100 |
| Q2 | 65% | 35% | $55 | $26 | 84 |
| Q3 | 45% | 55% | $51 | $23 | 66 |
| Q4 (2026) | 35% | 65% | $48 | $21 | 55 |
Three drivers behind the cost-per-demo column. Mix shift did the heavy lifting, since a human in a capture rig produces demonstrations 3 to 5 times faster than a teleop station. The QA feedback loop cut waste. And per-hour prices themselves drifted down as rigs amortized, consistent with the price trajectory table in our annual report. What did not drive it: operator pay cuts. Rates held flat; the savings were structural.
Teleop did not go away. It concentrated. By Q4 the customer bought fewer teleop hours but harder ones: long-horizon bimanual tasks and recovery-from-failure demonstrations, the categories where embodied action data has no substitute yet.
The Audit Nobody Scheduled
Midway through the program, the customer’s legal team requested proof of consent for every identifiable person appearing in the egocentric corpus, citing their EU AI Act readiness work (Regulation (EU) 2024/1689). Egocentric capture records bystanders by default, which makes this the hardest provenance problem in robotics data.
We got lucky in the boring way: our capture protocol had required signed consent keyed to episode IDs since before the program started. The audit became a five-day export and review rather than a retroactive consent hunt across 38,000 hours. I want to be precise about the counterfactual, because it is the real lesson: had those records not existed per episode, the corpus would have been unusable for their EU deployment plans, and no refund makes 38,000 hours of re-collection fast. Provenance is a property you build in at capture time or not at all. Our report predicts this becomes a standard procurement gate in 2027, and after this program we consider that prediction conservative.
What We Would Tell a Team Planning 2027
The transferable results are the structural ones, so here is the checklist we now bring to every program kickoff:
- Set a teleop floor, not a teleop default. Buy embodied hours for post-training and recovery behaviors; let video carry diversity.
- Sign usable-hour terms with a written rubric before the first batch. The rubric is cheap; the month-four dispute is not.
- Fund the rejection feedback loop. Weekly rejection reasons to operators halved our rejection rate in two months.
- Demand episode-keyed consent records now. Retroactive provenance does not exist.
- Re-run your mix ablation quarterly. This program’s biggest savings came from letting evidence overrule the original plan.
The customer’s evaluation success rate roughly doubled across the year, and their 2027 plan now starts where this one ended: video-heavy, usable-hour priced, curation funded first. The full market context, including our seven predictions and price benchmarks, is in The State of Robotics Training Data 2027.
Next Step
Next step: if you want the acceptance rubric template this program converged on, Book a Demo and we will walk through it against your task families, or Download Sample Data to see episode-keyed manifests firsthand.
Frequently Asked Questions
How much data does a VLA program actually buy in a year?
This program delivered roughly 2,100 usable teleoperation hours and 38,000 usable egocentric video hours over twelve months. The spend ratio moved from 80/20 teleop-heavy to 35/65 video-heavy as the customer’s ablations favored video scale.
What QA rejection rate should teams expect for teleoperation data?
This program started at 22 percent under raw-hour terms and stabilized at 10 to 12 percent after moving to usable-hour contracts with a written rubric and weekly operator feedback. DexSet’s broader benchmarks show 15 to 30 percent for programs without that loop.
Did shifting budget to egocentric video hurt model performance?
Not in this program. The customer’s 40-task evaluation success rate roughly doubled over the year while cost per successful demonstration fell about 45 percent. Model architecture also improved in parallel, so DexSet credits the data program for the cost curve, not the full model gain.
Why did consent provenance matter mid-program?
The customer’s legal team audited the egocentric corpus for EU AI Act readiness. Because consent records were keyed to episode IDs from day one, the audit took five days. Without per-episode records, 38,000 hours would have been unusable for EU deployment.
What would DexSet change in a 2027 program?
Start where this one ended: a teleop floor rather than a teleop default, usable-hour terms signed before the first batch, curation funded at 25 to 30 percent, and quarterly mix ablations that are allowed to overrule the original plan.
Enoch Pakanati
Enoch Pakanati is the strategic architect behind DexSet’s mission to become the undisputed market leader in robotics training data. He oversees the company’s growth strategy, focusing on capturing dominant market share across all data modalities required for modern robotics, including egocentric capture, teleoperation, and simulation-to-real data pipelines.
At DexSet, Enoch is responsible for transforming the company’s deep technical capabilities into a market-leading brand that foundation model labs and robotics OEMs trust implicitly. He focuses on scaling DexSet’s global footprint and ensuring the company stays ahead of the industry’s rapidly evolving data needs. His leadership is centered on one objective: making DexSet the singular, global standard for the data that powers the robotics revolution.