The State of Robotics Training Data 2027: DexSet’s Annual Report
In a 2027 budget review we joined last month, a Head of Data spent forty minutes defending a plan that was 70 percent teleoperation by spend, and the CFO ended the debate with one question: why are we paying collection rates for coverage the public corpora already give away? Nobody in the room had a clean answer. Versions of that meeting are happening across robot foundation model companies right now, because the model architecture is no longer the scarce ingredient. The scarce ingredient is the right data, at the right quality, at a price that survives exactly that kind of question. Our thesis for 2027, argued throughout this report: the binding constraint on robot learning programs is shifting from collecting hours to trusting them, and budgets that still optimize for volume will misallocate a meaningful share of their spend.
The answers from 2025 no longer hold. Teleoperation hours were the default purchase then. Egocentric human video was a research curiosity. Tactile sensing was something you read about at CoRL and quietly deferred. Pricing was quoted per raw hour, and nobody asked where the consent forms lived. All five of those assumptions are shifting at once.
This report is our attempt to map the shift. It covers where the field actually stands as of late 2026, backed by published research and public datasets, and then makes seven explicit predictions for 2027. The predictions are our opinion, formed from running capture rigs, QA pipelines, and pricing negotiations every week. We flag them as opinion so you can weigh them accordingly. The trend tables use our internal benchmarks, stated as ranges, because we would rather publish real numbers than hide behind “contact sales.”
DexSet supplies egocentric, exocentric, teleoperation, mono, and stereo data to physical AI teams, so we see demand shifts before they show up in survey reports. Read this the way you would read a peer’s planning memo: check our reasoning, steal the tables, and disagree where your evidence is better.
TL;DR: State of Robotics Training Data 2027 – VLA models are mainstream. RT-2, OpenVLA, π0, and GR00T made large-scale demonstration data the primary cost driver for robot learning programs. – Open X-Embodiment and DROID set the public baseline. Buying data that duplicates free public coverage is the most common budget mistake we see. – Our seven predictions for 2027, flagged as opinion: egocentric human video becomes the dominant pretraining modality; world models raise the video quality bar; tactile moves onto the spec sheet; usable-hour pricing replaces raw-hour pricing; consent provenance becomes a procurement gate under the EU AI Act; cross-embodiment transfer narrows the moat of robot-collected data; and quality metrics standardize. – Our benchmark pricing: bimanual teleop at $40 to $60 per usable hour in 2026, trending toward $34 to $55 in 2027. Egocentric human video at $18 to $34 per usable hour, trending toward $15 to $28. – Curation and QA will consume 25 to 30 percent of a well-run 2027 data budget, up from roughly 12 percent in 2025.
Where Robotics Training Data Stands Heading Into 2027
The state of robotics training data in late 2026 is defined by three settled facts: vision-language-action models are the mainstream architecture for manipulation, humanoid programs have moved from prototypes to fleet-scale data collection, and public datasets now set a quality and coverage floor that paid data must clear.
None of these was settled three years ago. RT-2 showed in 2023 that co-training on web-scale vision-language data and robot trajectories transfers semantic knowledge into control (Brohan et al., arXiv:2307.15818). OpenVLA made a 7B-parameter VLA openly available and reproducible in 2024 (Kim et al., arXiv:2406.09246). Physical Intelligence’s π0 demonstrated flow-matching action generation across many embodiments (Black et al., arXiv:2410.24164), and NVIDIA’s GR00T N1 pushed the humanoid foundation model framing into the open with a dual-system architecture trained on a pyramid of real robot data, human video, and synthetic data (arXiv:2503.14734). The architecture debate is not over, but the data implication is consistent across all four: these models are trained on demonstration data measured in thousands of hours, and demand grows with every scaling result.
The public baseline matters just as much. Open X-Embodiment aggregated over one million trajectories across 22 embodiments (arXiv:2310.08864). DROID added 76,000 teleoperated episodes collected across 564 scenes in 52 buildings (arXiv:2403.12945). On the human side, Ego4D contributed 3,670 hours of egocentric video (arXiv:2110.07058) and EgoExo4D paired egocentric and exocentric views of skilled activity with over 1,200 hours of footage (arXiv:2311.18259). Tooling matured alongside the data: Hugging Face’s LeRobot standardized dataset formats and training loops for the community (github.com/huggingface/lerobot).
The practical implication for buyers: paid data in 2027 has to be justified against what is free. If a vendor quotes you for generic tabletop pick-and-place episodes on a Franka arm, Open X-Embodiment already gives you a large volume of that for nothing. The paid budget belongs where public data is thin, which is exactly where the field is heading next.
The Modalities That Matter in 2027
A data modality, in robotics training terms, is a distinct capture format defined by viewpoint, sensor stack, and collection method: teleoperation trajectories, egocentric human video, exocentric multi-view video, mono or stereo streams, depth, IMU, gaze, tactile, and synthetic rollouts. Modality choice determines what a model can learn, and the demand mix across modalities is the single clearest signal of where the field is going.
Here is how we see budget share shifting, based on the RFPs and programs that cross our desk. The 2027 column is our projection, and it is opinion, not measurement.
| Modality | 2025 share of new data budgets | 2026 share (observed) | 2027 share (our projection) |
|---|---|---|---|
| Teleoperation (robot embodiment) | 55% | 45% | 30% |
| Egocentric human video (mono + stereo) | 15% | 25% | 40% |
| Exocentric multi-view | 10% | 10% | 10% |
| Simulation / synthetic | 15% | 12% | 8% |
| Tactile and force-torque annotated | under 1% | 3% | 8% |
| Curation, relabeling, QA of existing data | 4% | 5% | (tracked separately below) |
Three notes on reading that table. First, teleoperation is not dying; it is being repositioned as the high-value post-training layer while cheaper human video absorbs pretraining volume. The ALOHA line of work established how far low-cost bimanual teleop can go (Zhao et al., arXiv:2304.13705; Mobile ALOHA, arXiv:2401.02117), and that layer still needs real robot embodiment. Second, the egocentric number is the headline: it is the only modality whose share we project to grow by 15 points in one year. Third, simulation share shrinks in relative terms only; absolute synthetic volume keeps growing, but it is increasingly generated in-house rather than purchased.
Seven Predictions for 2027
A prediction, for the purposes of this report, is a falsifiable claim about the robotics data market twelve to eighteen months out. Everything in this section is Level 5 evidence on our own pyramid: DexSet opinion, informed by first-hand operations but not yet proven. We publish predictions so you can hold us to them in next year’s report.
Prediction 1: Egocentric human video becomes the dominant pretraining modality
Egocentric human video is first-person footage of people performing tasks, captured from head-mounted or glasses-form-factor rigs, often with stereo, IMU, and gaze. Our prediction is that by end of 2027 it accounts for the largest share of new pretraining data budgets at robot foundation model companies, ahead of teleoperation.
The reasoning is arithmetic plus published precedent. A human wearing a capture rig produces task demonstrations at 3 to 5 times the hourly rate of a teleoperated robot, at roughly half the fully loaded cost per usable hour on our benchmarks, which compounds into several times cheaper per finished demonstration. EgoExo4D showed the research community how to structure paired ego-exo skilled activity data at scale (arXiv:2311.18259), and GR00T N1 explicitly places human video in its data pyramid between web data and robot trajectories (arXiv:2503.14734). The open question is transfer efficiency, and that is exactly what cross-embodiment results keep improving. The teams we supply are already running 60/40 human-video-to-teleop ratios in new programs; in 2025 that ratio was inverted.
Prediction 2: World models raise the video data quality bar
A world model is a learned simulator that predicts future observations conditioned on actions, trained largely on video. Our prediction is that world model training becomes a major demand source for robotics video in 2027, and that it forces a quality step-change: stable exposure, calibrated stereo, consistent frame rates above 30 fps, and dense temporal coverage without cuts.
Policy learning tolerates compression artifacts and dropped frames better than next-frame prediction does. When your loss function is reconstruction of the future, every jitter and rolling-shutter wobble is signal you are teaching the model to hallucinate. In our own QA pipelines, footage that passes for imitation learning fails world-model acceptance criteria roughly a third of the time. Buyers should expect vendors to publish per-clip quality manifests, and vendors who cannot will lose world-model contracts.
Prediction 3: Tactile sensing goes from niche to spec-sheet requirement
Tactile data is contact-based sensing, from fingertip arrays and vision-based sensors to wrist force-torque, synchronized with visual streams and actions. Our prediction is that by late 2027, tactile channels appear as a standard line item in manipulation data RFPs, the way depth did around 2024, even though most buyers will still purchase it for only a subset of hours.
Contact-rich tasks are where current VLAs fail most visibly: insertion, cable routing, fabric handling, anything where vision alone is ambiguous at the contact point. The hardware is no longer exotic, and once two or three published results show clear gains from tactile-conditioned policies on those task families, procurement follows. We started offering force-torque and fingertip-array capture on our bimanual rigs in 2026; it is currently 3 percent of delivered hours and the fastest-growing line in our catalog.
Prediction 4: Usable-hour pricing replaces raw-hour pricing
A usable hour is sixty minutes of data that survives QA: correct task execution, complete sensor streams, valid calibration, and accurate labels. Our prediction is that by end of 2027, usable-hour pricing is the default commercial unit in robotics data contracts, and raw-hour quotes become a red flag.
Raw-hour pricing hides a 15 to 30 percent rejection rate that the buyer discovers only after delivery. We moved our own contracts to usable-hour terms because it aligned incentives: our operators get feedback from the same QA gate the customer sees. The market pressure is simple; once a few vendors quote usable hours with published acceptance criteria, every raw-hour quote looks like an undisclosed surcharge. Expect contracts to specify acceptance criteria as an appendix, the way SLAs specify uptime.
Prediction 5: Consent provenance becomes a procurement gate under the EU AI Act
Consent provenance is the documented chain showing that every person appearing in a dataset consented to capture and to the specific training use. Our prediction is that during 2027, provenance documentation shifts from a nice-to-have to a pass/fail gate in enterprise procurement, driven by the EU AI Act’s staged obligations.
The Act entered into force in August 2024 (Regulation (EU) 2024/1689), with general-purpose AI obligations applying from 2025 and high-risk system obligations phasing in through 2026 and 2027. Robots operating around people sit uncomfortably close to the high-risk categories, and data governance requirements flow down to training data suppliers. Egocentric capture makes this acute: a head-mounted camera records bystanders by default. Our capture protocol requires signed consent from every identifiable person and stores consent records keyed to episode IDs; in 2026 we started getting audited on it by customers’ legal teams for the first time. That is the leading edge of the gate.
Prediction 6: Cross-embodiment transfer narrows the moat of robot-collected data
Cross-embodiment transfer is a model’s ability to apply skills learned on one robot body to a different one. Our prediction is that continued transfer gains through 2027 reduce the premium buyers will pay for data collected on their exact robot, shifting value toward task and scene diversity instead of embodiment match.
Open X-Embodiment demonstrated that co-training across 22 embodiments improves per-robot performance (arXiv:2310.08864), and π0 and GR00T both train across heterogeneous fleets by design. If a policy pretrained on mixed embodiments fine-tunes to yours with 50 hours instead of 500, then the 450-hour difference stops justifying bespoke collection at premium rates. The data that retains pricing power is the data that is expensive to reach: rare scenes, long-horizon tasks, contact-rich manipulation, and human video at diversity levels no single robot fleet can match.
Prediction 7: Data quality metrics standardize
A data quality metric, in this context, is a quantitative, dataset-level measure a buyer can verify independently: calibration reprojection error, stream-synchronization tolerance, task success labeling accuracy, coverage entropy across scenes and objects. Our prediction is that by end of 2027 a de facto standard set of five to eight metrics appears in most serious RFPs, seeded by community formats like LeRobot’s dataset schema rather than by a standards body.
Today, every buyer invents their own acceptance spreadsheet and every vendor answers a different questionnaire. That is pure friction on both sides. The fix will look like what happened with model evaluation: nobody legislated benchmarks, but everyone converged on a shared handful because comparison demanded it. We publish our QA thresholds with every delivery for exactly this reason, and we would rather compete on measured quality than on adjectives.
Price Trajectory: Our Benchmarks Through 2027
Robotics data pricing is the fully loaded cost a buyer pays per usable hour of delivered, QA-passed data. The table below is built from DexSet’s own delivered contracts and rig cost models. These are our benchmarks, stated as typical ranges; 2027 figures are our projection and should be read as opinion.
| Data type | 2025 (per usable hour) | 2026 (per usable hour) | 2027 projection |
|---|---|---|---|
| Bimanual teleoperation (ALOHA-class rig) | $48 to $60 | $40 to $60 | $34 to $55 |
| Single-arm teleoperation | $30 to $58 | $28 to $52 | $28 to $45 |
| Egocentric human video, mono | $22 to $40 | $18 to $34 | $15 to $28 |
| Egocentric stereo with IMU and gaze | $28 to $40 | $24 to $38 | $18 to $34 |
| Exocentric multi-view (4+ cameras, calibrated) | $25 to $50 | $22 to $45 | $20 to $38 |
| Tactile-annotated manipulation (bimanual rate plus annotation and added QA) | rarely offered | $55 to $95 | $50 to $85 |
| Curation/QA of existing corpora (per hour reviewed) | $5 to $10 | $6 to $12 | $8 to $12 |
Two structural points. Collection prices fall because rigs amortize and operator pools deepen; the same forces that have pulled bimanual teleop rates down since 2025 will keep working. Curation prices rise because the work gets harder: world-model acceptance criteria, provenance audits, and standardized quality metrics all add review depth per hour. The crossing of those two curves is the quiet story of 2027 budgets.
The Economics: What a 2027 Data Budget Should Look Like
A robotics data budget is the annual allocation covering collection, curation, QA, licensing, and compliance for a model training program. On our benchmarks, a mid-size foundation model team spending $2M in 2027 should expect a very different split than the same $2M bought in 2025.
Using midpoint prices from the table above, and carving curation out first at 25 to 30 percent (call it $540K), $2M allocated on our projected 2027 mix buys roughly: 9,000 usable hours of bimanual teleop (about $400K), 25,000 usable hours of egocentric human video (about $560K at blended mono and stereo rates, with stereo on a third of hours), 4,500 hours of exocentric multi-view (about $130K), 1,700 hours of tactile-annotated manipulation (about $115K), with the remaining roughly $250K on licensing, compliance, and contingency. Teams that skip the curation carve-out and treat QA as free discover the cost anyway, in engineer-hours spent triaging bad episodes after training runs fail.
The ROI logic has also inverted. In 2025 the question was “how many hours can we afford.” In 2027 the question is “how few hours do we need at each layer of the pyramid.” Cross-embodiment pretraining plus human video means your expensive robot-hours target the last-mile gap, and every dollar moved from redundant collection to curation typically returns more model performance per dollar. That claim is our operational experience, not a published result, and we flag it as such.
Case Study: Twelve Months of Data for a Humanoid Foundation Model Team
This case study describes a year-long program DexSet ran for a humanoid foundation model team, anonymized at their request. It is Level 1 evidence: first-hand, but a single program, so generalize with care.
The team came to us in mid-2025 with a plan that was 80 percent teleoperation by budget. Over twelve months the program delivered roughly 2,100 usable teleop hours across bimanual and mobile-manipulation tasks and 38,000 usable hours of egocentric human video across kitchen, warehouse, and retail scenes. Three things changed along the way. First, their ratio flipped: by month eight, new orders were 65 percent human video by spend, because their pretraining ablations showed video scale moving downstream success more than marginal teleop hours. Second, rejection taught pricing: their first teleop batches saw a 22 percent QA rejection rate under raw-hour terms, which is what pushed both sides to usable-hour contracts with a published acceptance rubric. Third, compliance arrived mid-program: their counsel requested consent provenance for every identifiable person in the egocentric corpus, and because our consent records were keyed to episode IDs from day one, the audit took days, not months.
The outcome we can share: their fine-tuned policy’s success rate on a 40-task internal evaluation roughly doubled over the program, and their cost per successful task-demonstration fell by about 45 percent from the first quarter to the last. We cannot attribute the model gains to data alone; their architecture improved in parallel. But the cost curve is a data-program result, and it came from the mix shift and the QA gate, not from squeezing operator wages.
Plan Your 2027 Program: The DexSet Data Scorecard
A data program scorecard is a structured worksheet for scoring vendors and allocating budget across modalities before you sign anything. We have packaged the framework in this report as a downloadable scorecard: modality mix worksheet, usable-hour acceptance rubric, the seven quality metrics we expect to standardize, and a consent provenance checklist mapped to EU AI Act obligations.
If you are planning a 2027 budget, run your current plan through the scorecard first. If your teleop share is above 50 percent, your curation line is below 15 percent, or your vendor quotes raw hours, the report you just read explains why we would push back.
Related reading from this report series: Why Data Curation Is the 2027 Bottleneck, Comparing Modality Investments for 2027 Budgets, A Year-Long VLA Data Program Retrospective, and 5 Hidden Challenges in 2027 Data Planning.
Next Step
Download the 2027 Data Program Scorecard or Book a Demo to walk through your modality mix with our data operations team, or start smaller: Download Sample Data and run our egocentric stereo episodes through your own QA gate.
Frequently Asked Questions
What is the state of robotics training data heading into 2027?
VLA models like RT-2, OpenVLA, π0, and GR00T made large-scale demonstration data the main cost driver in robot learning. Public datasets (Open X-Embodiment, DROID, Ego4D, EgoExo4D) set the free baseline, while paid budgets shift toward egocentric human video, tactile sensing, and curation.
Which data modality will matter most for robot foundation models in 2027?
In our opinion, egocentric human video takes the largest share of new pretraining budgets in 2027, at roughly 40 percent, because it delivers task demonstrations at several times the rate and a fraction of the cost of teleoperation. Teleoperation remains essential for post-training on real robot embodiments.
How much does robotics training data cost in 2027?
On DexSet’s benchmarks, bimanual teleoperation runs $40 to $60 per usable hour in 2026, trending toward $34 to $55 in 2027. Egocentric human video runs $18 to $34 per usable hour, trending toward $15 to $28. Tactile-annotated manipulation is the premium tier at $50 to $85 projected.
What is usable-hour pricing in robotics data?
Usable-hour pricing charges only for data that passes an agreed QA gate: correct task execution, complete synchronized sensor streams, valid calibration, and accurate labels. It replaces raw-hour pricing, which hides rejection rates of 15 to 30 percent that buyers discover after delivery.
How does the EU AI Act affect robotics training data?
The EU AI Act (Regulation (EU) 2024/1689) phases in data governance obligations through 2027 that flow down to training data suppliers. For egocentric capture, which records bystanders by default, buyers increasingly require documented consent provenance keyed to individual episodes before purchase.
Does robot-collected data still have a moat if cross-embodiment transfer works?
A smaller one. Open X-Embodiment showed co-training across 22 embodiments improves per-robot performance, and π0 and GR00T train across heterogeneous fleets by design. Value shifts from embodiment match toward task diversity, rare scenes, contact-rich manipulation, and human video scale.
Are these predictions facts or opinions?
Opinions, and flagged as such throughout. The historical claims are cited to published research; the seven 2027 predictions and all forward-looking prices are DexSet’s judgment based on first-hand operations, published so readers can hold us accountable in next year’s report.
Sainath Gupta
Sainath Gupta is the visionary leader of DexSet, a company dedicated to building the foundational data layer for the robotics revolution. Sainath recognized early on that the primary bottleneck for physical AI wasn't hardware or model architecture, but the lack of high-quality, scalable training data.
At DexSet, he has pioneered a "manufacturing-first" approach to data collection. This includes the development of transparent cost-per-hour benchmarks and the scaling of global networks for egocentric and teleoperation capture. Sainath is a vocal advocate for the industry’s shift toward treating human experience as the primary pretraining substrate for robots. His leadership at DexSet is focused on one goal: providing the millions of hours of high-fidelity data required to bring humanoid robots out of the lab and into the real world.