Skip to main content

Dexset

Case Study: How We Scaled Data Quality & QA for Robotics Datasets for a VLA Model

Enoch Pakanati

A dataset audit is usually defined as a compliance exercise: verify the deliverable matches the invoice, file the report, move on. That definition hides what an audit actually is for a team training robot policies, which is the cheapest model improvement available to them, and this case study is the evidence for that claim. One corpus, one audit, a smaller dataset at the end of it, and a better model.

The engagement started with a problem that sounded like a training bug. A VLA team building a generalist manipulation model came to us with identical architecture, identical hyperparameters, and evaluation success that swung noticeably between checkpoints depending on which data deliveries were in the mix. Some batches helped. At least one, they suspected, actively hurt. They could not tell which without burning a GPU run per hypothesis.

The situation existed because they had done the rational-looking thing: bought roughly 1,000 hours of teleoperation data from three vendors plus their own internal rig to maximize diversity and delivery speed. Four sources meant four sets of conventions, four QA standards (or none), and no common audit before everything was pooled. Diversity without normalization had turned into variance.

This post walks through exactly what we did: the audit numbers, what failed and why, what the QA layer cost, and what changed in their training runs and their vendor contracts. Names and identifying task details are withheld by agreement; every number is real and rounded.

DexSet ran this engagement with the same three-tier pipeline we use on our own collection, which is documented in full in our guide to data quality and QA for robotics datasets.

Key Takeaways – A 1,000-hour mixed-vendor corpus yielded 74 percent usable hours after a full three-tier audit. – Tier 1 automated checks rejected 14 percent of hours; the single largest cause was an intermittent 2-frame action-observation desync inside one vendor’s delivery. – Human review flagged another 9 percent, dominated by success-label errors on near-miss episodes. – Retraining on the cleaned 740 hours produced more stable evaluation curves and higher held-out success than the raw 1,000 hours. – The audit cost roughly $9 per hour and paid for itself in renegotiated vendor terms alone.

The Starting Point: 1,000 Hours, Four Sources, Zero Common Standard

A mixed-source corpus is a training set assembled from multiple rigs or vendors without a shared validation standard, and it is the default state of most scaling teams’ data. The team’s corpus broke down as: internal rig (250 hours), Vendor A (350 hours), Vendor B (250 hours), Vendor C (150 hours). Tasks were tabletop and mobile manipulation, mostly bimanual, targeting an ACT-style architecture scaled toward VLA fine-tuning.

The symptoms will be familiar to anyone who has pooled data this way. Checkpoint-to-checkpoint evaluation variance far above what seed noise explains. A held-out task where success dropped after adding data. This is the single-team version of what Open X-Embodiment hit at field scale: heterogeneous robot data needs normalization and validation before pooling, or the heterogeneity itself becomes the dominant signal (arXiv:2310.08864).

Tier 1: Automated Checks Rejected 14 Percent

Automated signal checks are deterministic per-episode validations run on every stream, and on this corpus they rejected 140 of 1,000 hours. The breakdown:

CheckHours rejectedPrimary cause
Action-observation desync >1 frame58Vendor B’s rig: intermittent 2-frame (66 ms) camera lag from a buffering fault
Camera dropout / black frames37USB bandwidth saturation on internal rig’s third camera
Timestamp monotonicity18NTP re-sync events mid-episode at one vendor site
Gripper state mismatch12Worn actuator on one internal arm reporting stale aperture
Joint limit violations8Retargeting bug in Vendor C’s VR pipeline
Episode length outliers7Forgotten recordings, stalled operators

The Vendor B finding is the one to internalize. The buffering fault came and went, so 58 hours of a 250-hour delivery carried the lag, the affected episodes looked identical to the clean ones in playback, and the dataset was fully schema-valid in the LeRobot format the team used (github.com/huggingface/lerobot). A policy trained on those hours would have learned to act on observations 66 milliseconds stale. Given how imitation errors compound over a trajectory (the DAgger dynamic, Ross et al., arXiv:1011.0686), this batch was almost certainly the “actively hurts” delivery the team suspected.

Tier 2: Human Review Flagged Another 9 Percent

Human spot review is rubric-based scoring of sampled episodes by trained reviewers, and here it covered 20 percent of Tier 1 survivors plus everything flagged marginal. Extrapolated and then confirmed by targeted full review of suspect slices, it removed another 90 hours.

Success-label errors dominated: 60 of the 90 hours. The pattern was near-miss inflation. Episodes where the task technically completed but the robot displaced or toppled non-target objects were labeled successful, and one vendor’s annotators had marked grasp-and-drop episodes as successes when the drop happened after their labeling cutoff. Audited success-label accuracy across sources ranged from 91 to 97 percent; our shipping threshold is 95. The remaining 30 hours were strategy rejections: recovery flailing left unsegmented, and operator habits (like resting the arm against the table edge mid-task) that no team wants a policy to imitate. The ACT authors’ observation that human demos are noisy and stochastic is an architecture-level warning about exactly this material (arXiv:2304.13705).

Tier 3: Policy-in-the-Loop Caught the Invisible Batch

Policy-in-the-loop validation trains a probe policy per source batch and compares held-out success against a baseline, and it produced the audit’s most interesting result. Vendor C’s remaining hours passed Tiers 1 and 2 cleanly, yet their batch alone trained measurably below baseline. Manual inspection of trajectory statistics found the cause: operators had collected nearly every episode from a single base pose and approach direction. The batch was clean and narrow. Per-episode QA can never catch this; only training on it can.

The team kept the batch but reweighted it and required approach-pose diversity in the recollection order.

Results and What It Cost

The cleaned corpus came to 740 usable hours: a 74 percent usable-hour yield. Retrained on it, the team reported evaluation curves stable across checkpoints and held-out task success above the original 1,000-hour runs. Fewer hours, better model, which matches every internal ablation we have run.

ItemNumber
Raw corpus1,000 hours
Tier 1 rejections140 hours (14%)
Tier 2 rejections90 hours (9%)
Tier 3 action1 batch reweighted, recollection spec changed
Usable-hour yield74%
Audit cost~$9 per hour (~$9,000 total)
Per-source yield range62% (Vendor B) to 88% (internal rig)

What Changed After the Audit

The lasting value of a QA audit is the process changes it forces, not the hours it rescues. Four changes stuck with this team:

Pre-flight checks at every station. A 90-second automated routine (sync verification, black-frame test, gripper sweep against measured aperture) now runs before each collection shift, internal and vendor alike. The desync that cost Vendor B a fifth of their delivery would have been caught on day one.

Ingest gating instead of batch auditing. Episodes now pass Tier 1 within minutes of recording, so a failing rig generates a maintenance ticket the same shift rather than a rejection report the same quarter.

Yield and label accuracy as contract terms. New vendor agreements specify the check catalogue, thresholds, 95 percent audited success-label accuracy, and a usable-hour yield floor, with re-audit rights on samples. Pricing is now discussed per usable hour.

Diversity requirements in collection specs. Recollection orders specify minimum dispersion on base poses and approach directions, closing the clean-but-narrow gap that Tier 3 exposed.

The per-source yield table also became commercial ammunition. Vendor B’s effective price per usable hour was 60 percent above their invoice price, and the team renegotiated with both underperforming vendors using yield and audited label accuracy as contract terms. The $9,000 audit was recovered before the next purchase order.

Next Step

The full methodology behind this audit, including every check threshold and the metrics dashboard, is in the complete guide: Data Quality & QA for Robotics Datasets. If your training curves swing between deliveries, book a pipeline review and we will run a sample of your corpus through the same three tiers.

Frequently Asked Questions

What rejection rate should I expect when auditing a mixed-vendor robot data corpus?

Our production range is 10 to 30 percent of raw hours. This corpus landed at 26 percent rejected (74 percent yield), with per-source yields from 62 to 88 percent.

An intermittent 2-frame action-observation desync affecting 58 hours of one vendor’s delivery. It was invisible to human review and schema validation, and it teaches policies a systematically wrong observation-action mapping.

Roughly $9 per hour, about $9,000 for the 1,000-hour corpus, covering automated checks, 20 percent sampled human review, and per-source probe-policy training.

No. The team reported higher held-out task success and more stable evaluation curves training on 740 cleaned hours than on the raw 1,000, consistent with compounding-error dynamics in imitation learning.

Put usable-hour yield, the automated check catalogue with thresholds, and audited success-label accuracy (95 percent or better) into the contract, with re-audit rights on samples.

Enoch Pakanati
Written by

Enoch Pakanati

Enoch Pakanati is the strategic architect behind DexSet’s mission to become the undisputed market leader in robotics training data. He oversees the company’s growth strategy, focusing on capturing dominant market share across all data modalities required for modern robotics, including egocentric capture, teleoperation, and simulation-to-real data pipelines.

At DexSet, Enoch is responsible for transforming the company’s deep technical capabilities into a market-leading brand that foundation model labs and robotics OEMs trust implicitly. He focuses on scaling DexSet’s global footprint and ensuring the company stays ahead of the industry’s rapidly evolving data needs. His leadership is centered on one objective: making DexSet the singular, global standard for the data that powers the robotics revolution.