Two teams we worked alongside last year signed for near-identical manipulation datasets at headline rates within a few dollars of each other. One closed its program roughly on budget. The other overran by nearly half and cut a planned training run to pay for it. The difference was never the rate. It was five quieter line items that neither quote itemized, and only one team went looking for them before signing.
I review DexSet’s QA ledgers, so I watch this pattern from the inside, and it supports one claim, which is the thesis of this post: data budgets die in the gap between quoted raw hours and delivered usable hours, and the five costs below are that gap, itemized. Teams negotiate hard on the dollars-per-hour figure, sign, and then watch these quieter items add 30 to 80 percent to the program. By the time the overrun is visible, the training run is scheduled and there is no negotiating position left.
These five costs stay hidden for a structural reason: they live in the gaps between quote and delivery. A quote prices raw hours; delivery is measured in usable, annotated, retrievable hours. Everything between those two definitions is where money leaks, and vendors have little incentive to itemize a gap that flatters their pricing.
This post names the five, puts our real numbers on each, and gives you the fix. All figures come from DexSet’s own pipeline benchmarks and program ledgers; where public hardware anchors exist, like the roughly $20k ALOHA rig, they are cited.
Read it before you sign anything. Then take the checklist at the bottom into your next vendor call.
Key Takeaways
- QA rejection (10 to 30% in our pipelines) is the largest hidden cost: a $40 quote at 25% rejection is really $53.33 per usable hour.
- Annotation scope creep adds $8 to $25 per hour per pass; label only selected data, not the whole corpus.
- Rig downtime and recalibration silently cut utilization; every idle shift raises amortization per hour.
- Storage, egress, and versioning run 3 to 6% of capture spend and spike at training time.
- Operator turnover resets a 30 to 50% throughput learning curve; retention is a data-cost lever.
1. QA Rejection: The Gap Between Raw Hours and Usable Hours
QA rejection cost is the money spent capturing episodes that never enter your training set, and it is invisible in any quote expressed per raw hour. In our pipelines, 10 to 30 percent of raw episodes fail review: dropped frames, desynchronized views, occluded end-effectors, failed task completions.
The math bites harder than teams expect:
| Quoted rate | Rejection rate | Real cost per usable hour | Hidden premium |
|---|---|---|---|
| $30 | 10% | $33.33 | +11% |
| $40 | 25% | $53.33 | +33% |
| $50 | 30% | $71.43 | +43% |
The fix. Three contract clauses. One: the vendor reports a measured rejection rate from a comparable program, not an aspiration. Two: recollection of rejected episodes is priced in writing, ideally on the vendor’s account above an agreed threshold. Three: you run a 100-to-200-hour paid pilot scored against your acceptance spec before any volume commitment. On our programs, a versioned task spec alone typically pulls rejection from the high 20s to under 15 percent within a few weeks.
2. Annotation Scope Creep: Paying for Labels You Never Train On
Annotation scope creep is the gradual expansion of labeling passes across an entire corpus when only a subset of the data needs them. Each pass costs real money: $8 to $12 per data hour for language instructions, $10 to $16 for subtask segmentation, $18 to $25 for dense masks and contact labels, per our benchmarks.
The failure mode is ordering “full annotation” on day one, before anyone knows which slices the model will actually consume. A 10,000-hour corpus with three blanket passes at a mid-range $38 per hour combined is $380,000 of labels, and in our experience a meaningful fraction of densely labeled episodes never influence a training run.
The fix. Stage it. Label language instructions broadly if your VLA needs them, then gate expensive passes behind data selection: annotate the episodes your curriculum actually samples. Run labeling in tranches with a two-week lag behind training experiments, so label spend follows demonstrated need. Teams that stage annotation typically spend 40 to 60 percent less on labels for the same eval performance.
3. Rig Downtime and Recalibration: Utilization Is the Denominator
Downtime cost is the amortization you pay while a rig is not collecting: maintenance, recalibration, resets, and idle shifts all raise the hardware cost of every hour that does get captured. A roughly $20k ALOHA-class cell (https://arxiv.org/abs/2304.13705), or a $32k Mobile ALOHA (https://arxiv.org/abs/2401.02117), is cheap only when it runs.
The numbers move fast. At two-shift utilization over 18 months, amortization is $4 to $9 per hour. Single shift with 50 percent idle time, and the same rig charges you $15 or more per hour before anyone touches a leader arm. Multi-camera exocentric arrays are worse offenders: calibration after every scene change eats 5 to 15 percent of scheduled collection time if it is not engineered out.
The fix. Treat utilization as a weekly KPI. Schedule calibration and scene resets into shift handovers, keep spare grippers and cameras on the shelf (a $600 spare beats a lost shift), and pre-stage scenes so operators walk into ready cells. If you are buying rather than building, ask the vendor how many shifts their rigs run; their utilization sets the amortization share baked into your rate.
4. Storage, Egress, and Versioning: The Bill That Arrives at Training Time
Data infrastructure cost is the spend on storing, versioning, and moving your dataset, and it stays invisible until the first big training run pulls the whole corpus out of cloud storage. Multi-view stereo capture generates terabytes per week; a 10,000-hour multi-camera program can produce several hundred terabytes before compression decisions are made.
Our planning figure is 3 to 6 percent of capture spend for storage, format conversion, and dataset versioning, with egress as the spike risk: pulling a few hundred terabytes across clouds at list egress prices can add tens of thousands of dollars per full-corpus read.
The fix. Decide storage format and residency before collection starts, not after. Co-locate data with training compute to kill egress. Standardize on a training-ready format on delivery (for example, LeRobot-compatible datasets, https://github.com/huggingface/lerobot, rather than raw ROS bags), so you pay conversion once. Version at the episode level so experiments pull slices, not the whole corpus.
5. Operator Turnover: The Learning Curve You Pay For Twice
Operator turnover cost is the throughput and quality you lose when a trained teleoperator leaves and a new one restarts the learning curve. In our programs, operators improve 30 to 50 percent in episodes-per-shift over their first 200 hours, and their rejection rates fall in parallel. Every departure resets both curves.
This cost hides inside blended rates. A vendor churning operators quietly delivers you a workforce that is permanently early-curve: slower, more rejected episodes, same invoice. You will never see a line item for it.
The fix. Ask vendors for operator tenure and how many hours their median operator has logged. In-house, pay experienced operators above generic labor rates; the throughput math justifies it easily. And instrument per-operator metrics, episodes per shift and rejection rate, so coaching happens before quality drifts.
The Pre-Signature Checklist
A pre-signature checklist converts these five hidden costs into questions a vendor must answer in writing before you commit volume:
- ☐ Measured QA rejection rate on a comparable program, and who pays for recollection
- ☐ Itemized rate card: capture, QA, each annotation pass ($8 to $25/hr range), infrastructure
- ☐ Rig utilization (shifts per day) behind the amortization in the rate
- ☐ Delivery format, storage residency, and who pays egress
- ☐ Median operator tenure and hours logged
- ☐ 100-to-200-hour paid pilot scored against your acceptance spec
If a vendor stalls on more than one of these, the hidden costs are not hidden from them. They are hidden from you. The full rate benchmarks behind every number in this post are published in our robot training data costs and pricing guide.
Audit Your Next Quote Against These Five
Ready to scope your data program? Talk to our team.
Frequently Asked Questions
What is the single biggest hidden cost in robot training data?
QA rejection. At the 10 to 30 percent rejection rates we measure, a quoted raw-hour rate understates the true cost per usable hour by 11 to 43 percent. It is the first number to demand from any vendor.
How much does annotation add to robot data costs?
$8 to $25 per data hour per pass in our benchmarks: language instructions at the low end, dense masks and contact labels at the top. Staging annotation behind data selection typically cuts label spend 40 to 60 percent.
How much should I budget for storage and data infrastructure?
Plan 3 to 6 percent of capture spend, and engineer egress out by co-locating data with training compute. Multi-view stereo programs can reach hundreds of terabytes, so format and residency decisions belong before collection starts.
Why does operator experience matter to my data budget?
Throughput improves 30 to 50 percent over an operator’s first 200 hours and rejection falls in parallel. High-churn workforces deliver permanently early-curve performance at the same hourly rate.
How do I catch these costs before signing a contract?
Use the six-question checklist above: measured rejection rate, itemized rates, rig utilization, delivery format and egress liability, operator tenure, and a paid pilot against your spec.