Skip to main content

Dexset

Why the 2027 Robotics Data Bottleneck Won’t Be Collection. It Will Be Curation.

Enoch Pakanati

Collected is not curated. That four-word distinction will decide more 2027 training outcomes than any modality choice, and it sits behind a failure mode we now see monthly. A robotics team raises a round, budgets a large data purchase, and receives exactly what the contract said: 5,000 hours of demonstrations, delivered on time. Six weeks later their training runs plateau, an engineer starts spot-checking episodes, and the real accounting begins. Twenty percent of episodes have desynced camera streams. Ten percent show the operator recovering from mistakes that were never labeled as recoveries. Task success labels disagree with the video in a nontrivial slice.

This happens because the industry solved collection first. Teleop rigs got cheap after ALOHA published its bill of materials (Zhao et al., arXiv:2304.13705), egocentric capture hardware became a consumer product category, and operator marketplaces scaled. Collection capacity has grown several-fold in three years. Curation capacity, the reviewing, filtering, relabeling, and verifying, still scales mostly with careful humans and slowly improving tooling. When one side of a pipeline scales dramatically and the other does not, the slow side becomes the bottleneck. That is our thesis: curation, not collection, becomes the binding constraint on robot learning programs in 2027, and teams that fund it as a first-class line item will out-train teams that treat it as overhead.

This post lays out why curation becomes the binding constraint, what it does to budgets, and what a defensible curation plan looks like. It draws on the QA benchmarks behind our annual report, The State of Robotics Training Data 2027.

DexSet runs curation as a gated pipeline on every hour we deliver, so the numbers below are first-hand. Where we speculate about 2027, we say so.

Key Takeaways – Collection capacity has scaled several-fold since 2023; curation still scales with human review. The slow side becomes the 2027 bottleneck. – On our benchmarks, 15 to 30 percent of raw collected hours fail a serious QA gate. Raw-hour pricing hides this; usable-hour pricing exposes it. – World-model training tightens acceptance criteria further: footage that passes for imitation learning fails world-model QA about a third of the time in our pipeline. – A well-run 2027 budget puts 25 to 30 percent into curation and QA, up from roughly 12 percent in 2025. – The fix is a written acceptance rubric, per-episode quality manifests, and curation funded as a first-class line item, not a residual.

What Data Curation Means in Robot Learning

Data curation is the set of processes that turn collected sensor recordings into training-ready data: filtering failed episodes, verifying stream synchronization and calibration, correcting task and phase labels, deduplicating near-identical demonstrations, and documenting provenance. Collection produces hours; curation produces usable hours, and models only benefit from the second kind.

The distinction sounds pedantic until you price it. A bimanual teleop hour that costs $50 to collect and fails QA is not a $50 loss; it is a $50 loss plus its share of training compute plus the engineer time spent diagnosing why the policy twitches. Imitation learning copies what it sees. If 8 percent of your demonstrations include unlabeled operator fumbles, your policy learns to fumble 8 percent of the time, and no amount of additional bad data fixes that.

Why Collection Outpaced Curation

Collection scaled because its costs are hardware and labor, both of which respond to money quickly; curation lagged because its costs are judgment and tooling, which respond slowly. That asymmetry, not any single failure, is what creates the bottleneck.

Consider what happened on each side. On collection: ALOHA-class rigs dropped bimanual teleop hardware into the low tens of thousands of dollars, Mobile ALOHA extended it to whole-home tasks (arXiv:2401.02117), and DROID demonstrated distributed collection across 52 buildings with a standardized platform (arXiv:2403.12945). Public aggregation through Open X-Embodiment normalized the idea that volume is achievable (arXiv:2310.08864).

On curation: the community got better formats, notably LeRobot’s dataset schema (github.com/huggingface/lerobot), and some automated checks. But the hard parts resist automation. Deciding whether a demonstration is a good example of “load the dishwasher” requires task understanding. Catching a subtly miscalibrated stereo pair requires checks most teams have not written. Judging whether an operator’s recovery from a slip is valuable data or noise is a genuine research question. Our own pipeline automates about 60 percent of checks by volume; the remaining 40 percent is trained human review, and that fraction has not moved much in two years.

The Numbers: What Curation Actually Costs

Curation cost is best expressed as a percentage of total data budget and a per-hour review rate, and both are rising. The table below combines our delivered-contract benchmarks with our 2027 projection, which is opinion.

Metric (our benchmarks)202520262027 projection
Raw-hour QA rejection rate, teleop15 to 30%15 to 30%15 to 30%
Additional rejection under world-model criterian/a~30% of IL-passing footage25 to 35%
Curation review cost per hour reviewed$5 to $10$6 to $12$8 to $12
Curation share of well-run data budgets~12%~18%25 to 30%
Automated share of QA checks (our pipeline)~45%~60%~70%

Two things to notice. First, review cost per hour rises even as automation improves, because acceptance criteria are tightening faster than tooling: world models punish frame drops and exposure instability that imitation learning shrugs off, and consent provenance review under the EU AI Act (Regulation (EU) 2024/1689) adds a compliance pass that did not exist in 2024. Second, rejection rates are not falling. Collection quality improves, but so does the bar, and the two roughly cancel.

What a Defensible 2027 Curation Plan Looks Like

A defensible curation plan is a funded, written pipeline with acceptance criteria agreed before collection starts. Based on what works in our own operations, it has five parts:

  • A written acceptance rubric. Task success definition per task family, stream-sync tolerance (we use 10 ms), calibration reprojection error threshold, and label accuracy sampling plan. If it is not written, it will be renegotiated after delivery.
  • Usable-hour contracting. Pay for hours that pass the rubric. This converts curation from your cost into a shared incentive.
  • Per-episode quality manifests. Every episode ships with its QA results. This is what makes later audits, including compliance audits, take days instead of months.
  • Curation funded at 25 to 30 percent. Carved out first, not left as a residual after collection spend.
  • A feedback loop to collection. Rejection reasons go back to operators weekly. Our rejection rate on mature task families drops by roughly half within two months once this loop runs.

Teams that adopt this in 2026 will feel the 2027 bottleneck as a manageable cost line. Teams that do not will feel it as a training-run mystery.

Where This Fits in the Bigger 2027 Picture

Curation is one of several shifts converging on 2027 data budgets, alongside the move to egocentric human video, tactile sensing, and usable-hour pricing. The full analysis, including modality demand tables and our seven predictions, is in our annual report: The State of Robotics Training Data 2027.

Next Step

Next step: if you are planning a 2027 budget, read the report and pressure-test your curation line against our benchmarks. If you want to see what a per-episode quality manifest looks like in practice, Download Sample Data and inspect ours.

Frequently Asked Questions

What is the biggest bottleneck in robotics training data for 2027?

Curation. Collection capacity has scaled several-fold since 2023 through cheap teleop rigs and consumer egocentric hardware, while curation still depends on human review. On DexSet’s benchmarks, 15 to 30 percent of raw hours fail a serious QA gate.

DexSet recommends 25 to 30 percent of the total data budget in 2027, up from roughly 12 percent in 2025. Review costs run $8 to $12 per hour reviewed on our projected benchmarks.

World models are trained to predict future frames, so compression artifacts, dropped frames, and exposure instability become training signal rather than tolerable noise. About a third of footage that passes imitation learning QA fails world-model criteria in DexSet’s pipeline.

An hour that passes an agreed acceptance rubric: correct task execution, complete synchronized sensor streams, valid calibration, and accurate labels. Usable-hour pricing exposes the rejection rate that raw-hour pricing hides.

Partially. DexSet automates about 60 percent of QA checks by volume (sync, calibration, completeness). Task-success judgment, label verification, and edge-case triage still require trained human review.

Enoch Pakanati
Written by

Enoch Pakanati

Enoch Pakanati is the strategic architect behind DexSet’s mission to become the undisputed market leader in robotics training data. He oversees the company’s growth strategy, focusing on capturing dominant market share across all data modalities required for modern robotics, including egocentric capture, teleoperation, and simulation-to-real data pipelines.

At DexSet, Enoch is responsible for transforming the company’s deep technical capabilities into a market-leading brand that foundation model labs and robotics OEMs trust implicitly. He focuses on scaling DexSet’s global footprint and ensuring the company stays ahead of the industry’s rapidly evolving data needs. His leadership is centered on one objective: making DexSet the singular, global standard for the data that powers the robotics revolution.