Skip to main content

Dexset

Four Ways to Buy Robot Training Data, Compared: Pros, Cons, and Costs

Taped above a monitor in a buyer’s office we visited last year was a coffee-stained, three-page acceptance spec, its 10 ms sync tolerance circled twice in red pen. That battered document was doing more procurement work than the forty-page RFP folder on the shelf beside it, because it forced every vendor quote onto the same axes. It also marked its owners as unusual: teams rarely frame “how are we going to buy this” as a decision at all. They email two vendors someone met at CoRL, pick the cheaper quote, and only discover they chose a procurement approach when it fails.

The failure is predictable because each buying approach has a known cost structure and a known blind spot. An informal purchase is fast and blind. A full RFP is thorough and slow. A pilot-first approach measures what matters but covers one vendor at a time. Open datasets are free and almost never match your embodiment or task distribution.

This post compares the four approaches on speed, cost, and risk, with the math that lets you pick deliberately. It draws on the same scorecard and pilot protocol as our full Robotics Data Buyer’s Playbook, which is where the reusable templates live.

At DexSet we respond to all four buying styles weekly, so we see their outcomes from the supplier side: which approaches produce clean contracts and which produce disputes about what “an hour of data” was supposed to mean.

Key Takeaways

  • There are four common procurement approaches: informal purchase, full RFP + scorecard, pilot-first, and open-data-plus-top-up. Each has a distinct cost and risk profile.
  • Informal buying is cheapest to run ($0 process cost) and most expensive to survive: yield surprises routinely add 20 to 40 percent to effective cost.
  • The RFP + scorecard + pilot combination costs roughly 3 to 5 weeks and $1,500 to $4,500 in pilot fees, and is the only approach that measures quality before annual commitment.
  • Open datasets (Open X-Embodiment, DROID, Ego4D) are excellent for pretraining and mixing, but embodiment and task mismatch means most teams still purchase targeted data on top.

Approach 1: Informal Purchase

An informal purchase is vendor selection without a written specification, scoring method, or pilot: the buyer requests quotes, reviews sample clips, and signs with the most convincing option. It is how most first data purchases happen, and it is defensible exactly once, at very small volume, when you are still learning what to specify.

Pros: fastest path to first data (days, not weeks); no process overhead; fine for exploratory volumes under 20 hours.

Cons: samples are curated, so latent defects (sync offsets, calibration drift) go undetected; quotes are not comparable because no shared spec exists; no contractual yield commitment, so failed hours are your loss; format surprises arrive with the first delivery.

Cost profile: zero process cost up front. In our benchmarks, yield surprises and conversion work typically add 20 to 40 percent to effective per-usable-hour cost versus a piloted vendor. At 1,000 hours, that is $8,000 to $16,000 of avoidable spend on a $40/hr program.

Approach 2: Full RFP with Weighted Scorecard

A full RFP approach sends a written acceptance spec and a fixed question set to five to eight vendors, then scores responses on weighted criteria before any commitment. This is classic procurement discipline adapted to robot data: quality SLAs, calibration and sync specs, throughput evidence, pricing transparency, licensing terms.

Pros: quotes become comparable because everyone bids the same spec; weak vendors self-eliminate (in our experience, roughly a third of recipients answer with adjectives instead of numbers); the scorecard creates an audit trail for the decision; licensing and consent problems surface before signature.

Cons: takes two to four weeks; still paper-based, so a vendor can score well and underdeliver; overkill below roughly $25,000 in annual data spend.

Cost profile: the process costs internal time only, typically 20 to 30 person-hours across spec writing, scoring, and reconciliation. It buys you comparability and eliminates the worst outcomes, but on its own it does not measure production quality.

Approach 3: Pilot-First

A pilot-first approach skips broad solicitation and goes straight to a 50-hour paid pilot with one or two candidate vendors, judged on predefined metrics: usable-hour yield, annotation audit accuracy, policy success delta on a fixed eval set, and loader time into LeRobot or RLDS. It optimizes for measured evidence over paper promises.

Pros: measures the only thing that matters, production output; small, fixed downside ($1,500 to $4,500 per pilot at market rates); fast when you already know the credible vendors; the policy-delta test catches defects no document review can.

Cons: covers only the vendors you pilot, so a better option may never be evaluated; sequential pilots take longer than parallel paper scoring; requires you to have a stable eval task set and baseline policy, which very early teams may lack.

Cost profile: $3,000 to $9,000 to pilot two vendors, plus about one engineer-week for evaluation. Expensive compared to reading PDFs, cheap compared to one bad quarter of deliveries.

Approach 4: Open Data Plus Targeted Top-Up

The open-data approach builds the base training mix from public corpora, then purchases only the targeted data the public sets cannot provide. The public layer is genuinely strong now: Open X-Embodiment spans over one million episodes across 22 embodiments (arxiv.org/abs/2310.08864), DROID adds 76,000 diverse teleop episodes (arxiv.org/abs/2403.12945), and Ego4D provides thousands of hours of egocentric human video (arxiv.org/abs/2110.07058), most of it accessible through Hugging Face dataset cards and RLDS tooling.

Pros: near-zero acquisition cost for pretraining scale; well-documented formats (RLDS, LeRobot conversions); community-validated quality.

Cons: embodiment mismatch (your gripper, camera placement, and control rates differ from the source robots); task distribution rarely matches your product; licenses vary and some restrict commercial use, so legal review is not optional; fine-tuning still demands in-domain demonstrations, which puts you back in one of the first three approaches for the data that moves your metrics most.

Cost profile: storage and engineering only for the public layer, then standard market rates ($28 to $60 per teleop hour, $15 to $40 per egocentric hour in our benchmarks) for the top-up volume, which is typically 10 to 30 percent of total hours but drives most of the task-specific performance.

Side-by-Side Comparison

The four approaches differ most in where they spend money: process time up front, or rework after delivery.

Approach Time to contract Process cost Quality measured before commitment? Typical effective cost penalty vs piloted baseline Best for
Informal purchase 3–10 days ~$0 No +20–40% Exploratory buys under 20 hours
Full RFP + scorecard 2–4 weeks 20–30 person-hours Partially (paper only) +5–15% Annual spend above $25k, multiple candidate vendors
Pilot-first 2–3 weeks per vendor $1,500–$4,500 per pilot Yes Baseline Teams with stable eval tasks and known vendor shortlist
Open data + top-up 1–2 weeks (legal + integration) Engineering time Yes for public layer, no for top-up unless piloted Depends on top-up approach Pretraining scale plus targeted fine-tuning

The pattern most mature buyers converge on is a hybrid: RFP to filter the field, scorecard to rank it, pilot to verify the winner, open data underneath it all as the pretraining base. That sequence is exactly what the Robotics Data Buyer’s Playbook packages, including the scorecard weights and pilot pass/fail thresholds.

One sequencing note from the supplier side: run the approaches in that order, not in parallel. Teams that pilot before writing a spec end up measuring vendors against criteria invented after the data arrived, which makes the results unarguable in exactly the wrong way; nobody can agree what a pass looks like. Teams that RFP without a spec get six incomparable quotes and mistake the spread for market variance. The spec is upstream of everything, takes about a week to write, and is the only artifact in the process that costs nothing but attention.

Next Step

Ready to scope your data program? Talk to our team.

Frequently Asked Questions

Which robot data procurement approach is cheapest overall?

The RFP-plus-pilot hybrid, once volume passes roughly $25,000 per year. Informal buying has the lowest process cost but the highest effective cost, because yield surprises add 20 to 40 percent on typical programs.

Not fully. Open X-Embodiment, DROID, and Ego4D are strong pretraining bases, but embodiment and task mismatch means fine-tuning still needs in-domain demonstrations, typically 10 to 30 percent of total hours purchased to spec.

For exploratory volumes under about 20 hours, where the goal is learning what to specify rather than feeding a production training run. Anything feeding a release model deserves at least a pilot.

Send the RFP to five to eight vendors, score all responses, then pilot the top one or two. Piloting more than two rarely changes the decision and doubles the evaluation load.

Four numbers agreed in advance: usable-hour yield (target 85 percent or higher), annotation accuracy on an independently re-labeled 5 percent sample (97 percent or higher), policy success delta on a fixed eval set, and loader time into your training format (one engineer-day or less).

Teleop, Egocentric, Exocentric, or Handheld: Comparing Robot Training Data Approaches by Cost

In February, we priced the same 5,000-hour manipulation corpus three different ways for one buyer: all bimanual teleoperation at a blended $44 per hour, a 65/35 egocentric-to-teleop mix at just under $30, and a handheld-gripper-heavy plan in between. Same task list, same acceptance spec. The spread between the first two plans came to roughly $70,000.

I build teleop cells for a living, so this is not an argument against teleoperation. It is the thesis that quoting exercise made unavoidable: modality selection is a budgeting decision, and the right modality for each training objective is the cheapest one that actually satisfies it. For plenty of what your model needs to learn, the premium teleop hour is simply the wrong purchase.

Teams get this wrong in both directions. Some buy 20,000 hours of premium teleop and burn budget teaching their encoder what a kitchen looks like, a job $18-per-hour egocentric video does fine. Others go all-in on cheap human video and then discover their policy has beautiful representations and no idea how to move an actual gripper.

The reason the mistake is so common is that modality costs and modality capabilities are usually discussed separately. Cost tables live in procurement decks; capability arguments live in arXiv papers. This post puts them in one place, with DexSet’s operating benchmarks attached, so you can match each dollar to the learning objective it actually serves.

By the end you will have per-hour costs for four collection approaches, an honest pros-and-cons list for each, and a decision matrix that maps training objectives to the cheapest modality that satisfies them.

Key Takeaways

  • Teleoperation ($28 to $60/hr) is the only approach that outputs executable robot actions natively. Pay for it where action supervision matters.
  • Egocentric human video ($15 to $40/hr) is the cheapest volume play, best for representation pretraining, worst for the embodiment gap.
  • Multi-view exocentric capture ($20 to $50/hr) buys scene context and cross-view consistency; calibration labor is its hidden cost.
  • UMI-style handheld grippers (under $1k per device by our estimates) collect gripper-centric data at near-egocentric labor cost, with heavier post-processing.
  • QA rejection (10 to 30 percent in our pipelines) and annotation ($8 to $25/hr per pass) apply to all four. Compare on cost per usable hour.

The Four Approaches, Defined

A data collection approach is the pairing of a capture device with a control source: robot teleoperation, head-mounted egocentric capture, calibrated exocentric camera arrays, or handheld instrumented grippers. Everything else, mono versus stereo, camera count, annotation depth, is a variation within these four.

Teleoperation: $28 to $60 per Hour

Teleoperation data is produced by a human driving a real robot through leader-follower arms or a VR interface, so every recorded frame pairs observations with executable actions in the robot’s own action space. The ALOHA project showed this could be done on a roughly $20k bimanual rig (https://arxiv.org/abs/2304.13705), and Mobile ALOHA extended it to whole-body mobile tasks on roughly $32k of hardware (https://arxiv.org/abs/2401.02117).

Pros

  • Native action labels; feeds imitation learning and VLA post-training directly
  • Matches your exact embodiment, gripper, and camera placement
  • Long-horizon, contact-rich tasks are demonstrable at production quality

Cons

  • Highest labor cost: trained operators at $18 to $38 per hour, plus rig amortization and QA
  • Throughput capped by operator skill and episode reset time
  • Data is embodiment-specific; switching robots strands some of its value

Egocentric Human Video: $15 to $40 per Hour

Egocentric data is first-person video from a head-mounted camera while a human performs tasks with their own hands, capturing human-level dexterity with no robot in the loop. Ego4D and EgoExo4D (https://arxiv.org/abs/2311.18259) made the research case; the commercial case is pure economics, since the collector works at natural speed on tasks they already know.

Pros

  • Cheapest per hour; scales to thousands of hours quickly
  • Enormous task and scene diversity, including real homes
  • Strong pretraining signal for visual encoders and hand-object interaction priors

Cons

  • No robot actions; the embodiment gap means it rarely supervises control directly
  • QA rejection skews high (motion blur, gaze drift, occlusion), 15 to 30 percent in our pipelines
  • Needs retargeting or paired data to transfer to a gripper

Multi-View Exocentric Capture: $20 to $50 per Hour

Exocentric data is third-person video from multiple calibrated, synchronized cameras observing the same task, giving models scene-level context and cross-view consistency that neither ego nor teleop streams provide alone. Stereo pairs add 15 to 25 percent over mono at the same view count and buy metric depth in return.

Pros

  • Full-scene coverage; occlusions in one view are recovered in another
  • Calibrated multi-view supports 3D reconstruction and world-model training
  • Pairs well with egocentric streams (the EgoExo4D recipe)

Cons

  • Calibration and synchronization labor at every scene change is the silent budget eater
  • Fixed arrays limit scene diversity; mobile arrays raise cost
  • Still no action labels without a paired control source

Handheld Instrumented Grippers (UMI-style): Near-Egocentric Cost, Gripper-Centric Output

Handheld gripper capture uses a portable, wrist-camera-equipped gripper operated by a human, producing gripper-centric trajectories without any robot present at collection time. The UMI paper (https://arxiv.org/abs/2402.10329) defined the category; our build estimate is under $1,000 per device including the camera.

Pros

  • Capex is trivial next to a $20k to $32k teleop cell
  • Collection happens anywhere a person can walk, at near-egocentric labor rates
  • Output is closer to robot action space than raw human video

Cons

  • Heavier post-processing to recover clean actions (SLAM drift, kinematic mismatch)
  • Gripper form factor constrains which tasks are demonstrable
  • QA tooling for this modality is younger; expect iteration

Master Comparison Table

Approach DexSet cost/raw hr Capex per station Action labels QA rejection Best use
Teleoperation $28 to $60 $20k to $32k Native 10 to 25% VLA post-training, imitation learning
Egocentric video $15 to $40 $300 to $3.5k None 15 to 30% Encoder pretraining, dexterity priors
Exocentric multi-view $20 to $50 $5k to $15k (3 to 8 cams) None 10 to 20% Scene context, world models, 3D
Handheld gripper (UMI-style) $18 to $42 (our estimate) Under $1k/device Recoverable 15 to 25% Diverse-scene manipulation at low capex

Annotation is additive to every row: $8 to $12 per hour for language instructions, up to $18 to $25 for dense masks and contact labels. And every row’s real price is its cost per usable hour: divide by (1 minus rejection rate). The full math, with a budget spreadsheet, lives in our robot training data costs and pricing guide.

Decision Matrix: Match the Dollar to the Objective

A modality decision matrix assigns each training objective the cheapest approach that actually satisfies it, instead of defaulting everything to the premium modality. Here is the one we use in scoping calls:

Your objective Buy this Not this Why
Post-train a VLA on your robot Teleoperation Egocentric You need native actions on your embodiment
Pretrain visual encoders at volume Egocentric Teleoperation Paying $42/hr for pixels is waste
Scene diversity across 100+ homes Handheld gripper or egocentric Fixed exo array Portability beats calibration
Depth-dependent manipulation Stereo exo + teleop Mono anything Metric depth earns its 15 to 25% premium
World-model or video-prediction training Exo multi-view + ego pairs Teleop only Cross-view consistency is the signal
Bimanual, contact-rich skills Teleoperation (ALOHA-class) Handheld gripper Two grippers, force-aware demos

The pattern behind the matrix: mix modalities and stage them. Open X-Embodiment’s 1M+ trajectories across 22 embodiments (https://arxiv.org/abs/2310.08864) and DROID’s 76k episodes (https://arxiv.org/abs/2403.12945) already prove cross-source data mixes train better generalists. Your budget should look like a portfolio, not a single line item. A 70/30 split of cheap pretraining hours to teleop post-training hours routinely cuts blended cost by a third in programs we run, with no loss on the action-supervised objectives.

Match the Modality to the Objective

Ready to scope your data program? Talk to our team.

Frequently Asked Questions

What is the cheapest way to collect robot training data?

Egocentric human video, at $15 to $40 per hour in our benchmarks, since hardware is a wearable camera and collectors work at natural speed. It is cheapest per hour but supplies no robot actions, so it cannot carry a program alone.

When you need executable actions on your specific embodiment: VLA post-training, imitation learning for contact-rich or bimanual skills. For those objectives nothing cheaper substitutes, which is exactly why you should not spend teleop dollars on anything else.

They are production-useful with caveats. Capex is under $1,000 per device by our estimates and scene diversity is the widest of the four approaches, but plan for heavier post-processing and a 15 to 25 percent rejection rate while the QA tooling matures.

Stereo adds 15 to 25 percent to exocentric capture cost and pays for itself on depth-dependent manipulation tasks. For pure representation pretraining, mono volume usually beats stereo precision per dollar.

Convert everything to cost per usable hour: quoted rate divided by (1 minus the measured QA rejection rate), plus annotation per pass. Our pricing guide includes the worked tables.

Start with the Robot Training Data Costs and Pricing Guide for the full benchmark tables, then book a scoping call and we will run your task list through the decision matrix above.

Why Training Data Costs Are the Biggest Bottleneck in Physical AI

“Training data” is usually defined as something you gather. That definition hides the economic fact that decides robotics budgets: robot demonstration data cannot be gathered at all. Language models scraped trillions of tokens the internet had already produced for free. A robot demonstration has to be manufactured, one episode at a time, by a person and a machine in a room, and manufactured goods have unit costs that scraping never did.

That manufacturing has a price, and the price is the thesis of this post: the binding constraint on physical AI progress is the unit economics of demonstration data, not model architecture. Our benchmarks put teleoperation at $28 to $60 per hour, egocentric human video at $15 to $40, and multi-view exocentric capture at $20 to $50, before annotation adds another $8 to $25 per pass. Multiply any of those by the hundreds of thousands of hours that scaling curves suggest, and the number stops looking like a data budget and starts looking like a Series B.

Compute costs fall on a curve you can plan around. GPU-hours get cheaper every year; teleoperator-hours do not, because they are wages plus hardware plus QA. So while everyone argues about architectures, the teams actually shipping robot foundation models are constrained by a much less glamorous question: how many usable demonstration hours can we afford this quarter?

This post breaks down why the bottleneck is economic rather than algorithmic, what the per-hour math actually looks like, and where the cost curve is bending. It draws on DexSet’s own capture operations, so the numbers are operating benchmarks, not estimates.

Key Takeaways

  • Robot data is manufactured, not scraped. Teleop costs $28 to $60 per hour; egocentric video $15 to $40; multi-view exo $20 to $50 (DexSet benchmarks).
  • QA rejection of 10 to 30 percent inflates every quoted rate. Budget on cost per usable hour.
  • Open X-Embodiment needed 21 institutions to pool 1M+ trajectories across 22 embodiments; no single lab could afford that collection alone.
  • Rig capex is the small part: an ALOHA-class station is roughly $20k and amortizes fast. Labor and QA dominate.
  • The cost curve bends through cheaper capture devices (UMI-style grippers), human video pretraining, and better data selection, not through cheaper wages.

The Bottleneck Is Economic, Not Algorithmic

The physical AI bottleneck is the gap between the demonstration volume that current methods need and the demonstration volume that current budgets can buy. Imitation learning works. ACT on ALOHA hardware showed fine bimanual manipulation from a rig that cost roughly $20k (Zhao et al., https://arxiv.org/abs/2304.13705). VLA models generalize further as data grows. The recipe is not the mystery; funding the recipe is.

Look at what it took to build the field’s reference datasets. Open X-Embodiment pooled data from 21 institutions to reach more than 1 million trajectories across 22 robot embodiments (https://arxiv.org/abs/2310.08864). DROID took a multi-lab consortium collecting across 52 buildings for a year to produce 76,000 episodes (https://arxiv.org/abs/2403.12945). These are consortium projects because the economics forced them to be. When the leading academic labs in the world have to carpool, the per-hour cost of data is the constraint worth studying.

Contrast that with a startup’s position. A humanoid company that wants 50,000 proprietary teleop hours at a blended $42 per hour is staring at a $2.1M capture bill before annotation, before storage, and before the 10 to 30 percent QA rejection rate we measure in our own pipelines pushes the real figure higher. That is the bottleneck in one sentence: the marginal trajectory costs real money, and scaling laws demand a lot of margins.

Where the Money Actually Goes

A fully loaded data cost is the sum of hardware amortization, collection labor, QA review, annotation, and infrastructure, and its composition explains why the bottleneck resists quick fixes. Hardware is the layer everyone obsesses over and the one that matters least.

Cost layer Teleoperation Egocentric video Share of total (typical)
Hardware amortization $4 to $9/hr $1 to $4/hr 10 to 15%
Collection labor $18 to $38/hr $10 to $26/hr 55 to 65%
QA and recollection $6 to $12/hr $5 to $10/hr 20 to 30%
Total (raw hour) $28 to $60/hr $15 to $40/hr 100%

Two things jump out of that table. Labor dominates, and labor does not follow Moore’s law. A teleoperator in year three costs what a teleoperator cost in year one, adjusted upward for wages. The only labor lever is throughput: in our programs, operator productivity improves 30 to 50 percent over their first 200 hours, which is real but bounded.

The second thing: QA is a fifth to a third of the bill, and it is the layer buyers most often forget. An episode fails for dropped frames, a desynced camera, an occluded gripper, or a task that did not actually complete. At a 25 percent rejection rate, a $40 quote is really $53.33 per usable hour. We walk through that math, with tables, in our full robot training data costs and pricing guide.

Why Compute Got Cheap and Data Did Not

Compute costs fall because silicon improves and utilization tooling matures, while demonstration data costs stay flat because their main input is human time in physical space. This asymmetry is the strategic fact of the next five years of robotics.

A training run you could not afford in 2023 is routine in 2026. But the demonstration hour you collected in 2023 cost about what it costs today, and the scene setup, the resets between episodes, and the review pass all still happen at human speed. Physics does not batch. You cannot checkpoint a kitchen.

The practical consequence: data spend is becoming the durable moat while compute spend becomes a commodity line item. Teams that treat their data budget with the same rigor as their compute budget, tracking cost per usable hour, rejection rates, and hours-to-policy-improvement, compound an advantage that a bigger cluster cannot erase.

Where the Cost Curve Actually Bends

Cost-curve bending in robot data comes from cheaper capture devices, cheaper modalities for pretraining, and better data selection, not from paying people less. Three developments are doing real work right now.

  • Handheld capture devices. UMI-style grippers (Chi et al., https://arxiv.org/abs/2402.10329) put a wrist camera on a portable gripper, so collection happens in real homes without a robot present. Our build estimate is under $1,000 per unit. Labor cost drops toward egocentric rates while output stays gripper-centric.
  • Human video for pretraining. Egocentric data at $15 to $40 per hour can carry representation learning, reserving expensive teleop for post-training. A 70/30 ego-to-teleop mix can cut blended cost per hour by a third without giving up action supervision where it counts.
  • Data selection over data volume. Deduplication, difficulty-aware sampling, and rejecting low-information episodes before annotation mean you pay $8 to $25 per hour of labels only on data that earns it.

None of these eliminate the bottleneck. They move the ratio of usable hours per dollar, which is the correct objective.

What This Means for Your Budget

A defensible data plan starts from cost per usable hour and works backward to volume, rather than starting from a raw-hour quote and hoping. If you take one action from this post, make it this checklist:

  • Get every vendor quote itemized across hardware, labor, QA, annotation, and infrastructure.
  • Demand a measured QA rejection rate from a comparable program, and pilot before committing volume.
  • Split your pipeline: cheap modalities for pretraining volume, teleop for action-supervised post-training.
  • Track cost per usable hour monthly. It is your burn rate’s most honest line.

Put the Numbers to Work

Ready to scope your data program? Talk to our team.

Frequently Asked Questions

Why is robot training data more expensive than language data?

Language data was scraped from the internet at near-zero marginal cost. Robot data is manufactured: a person, a rig, and a physical scene produce one episode at a time, at $15 to $60 per hour depending on modality, plus QA and annotation.

Collection labor, at 55 to 65 percent of the fully loaded hourly rate in our programs. Hardware amortization is only 10 to 15 percent, which is why cheaper rigs alone do not fix the bottleneck.

As a reference point, 50,000 teleop hours at a blended $42 per hour is $2.1M before annotation and before QA rejection losses of 10 to 30 percent. Consortium datasets like Open X-Embodiment exist precisely because no single lab wanted to carry that cost.

It reduces it for some skills, but sim-to-real transfer still needs real-world demonstrations for contact-rich manipulation, and mixed pipelines still budget significant real capture. Treat sim as a multiplier on real data, not a replacement.

Tighten task specs to cut QA rejections, mix cheaper egocentric or UMI-style capture into pretraining, annotate only selected data, and measure rejection rates continuously. The full cost model is in our pricing guide.

The Robot Training Data Costs and Pricing Guide publishes our complete per-hour benchmarks, rig capex table, and a downloadable budget spreadsheet. No sales call required to see the numbers.

Why Exocentric & Multi-View Data Is the Biggest Bottleneck in Physical AI

Language models got their training data for free: by the time the first large transformer trained, the internet had already spent decades producing trillions of tokens describing every topic from millions of viewpoints. Physical AI inherits no such gift. Robot data must be manufactured episode by episode, and most of it has been manufactured the cheap way: one camera, no calibration record, no synchronization guarantees. Compute got cheaper, architectures converged, and the gap between what models could absorb and what capture pipelines produce kept widening.

The pattern shows up concretely in teams we work with. One VLA group spent six weeks tuning architectures against a plateau. Bigger backbone, longer action chunks, better augmentation. The success rate on cluttered scenes moved two points. Then they looked at their data and found the real ceiling: every one of their 90,000 episodes was recorded from a single camera, and 40 percent of failures happened in the exact frames where that camera could not see the target object.

This post makes the case that calibrated exocentric and multi-view capture, not model design, is the binding constraint on manipulation performance right now, and shows what closing the gap costs. The evidence comes from public datasets (DROID, Open X-Embodiment, Ego-Exo4D), the RoboMimic observation-space study, and our own capture benchmarks at DexSet, where multi-view teleoperation rigs are what we run every day.

Key Takeaways – Single-view episodes cap policy performance on occlusion-heavy tasks regardless of architecture; the failure is in the data, not the model. – DROID made two external stereo views plus wrist a hard protocol requirement across 76,000 episodes; Open X-Embodiment’s viewpoint chaos shows what happens without such a standard. – The bottleneck is operational, not scientific: calibration drift, sync skew, and 3-4x storage are why teams default to one camera. – Our benchmarks: DROID-style rigs cost $4,500 to $7,000 to build and $26 to $38 per operated capture hour; that premium is small against a wasted training run.

Why Single-View Data Caps Policy Performance

A single-view dataset gives the policy exactly one projection of the world per timestep, so any state the camera cannot resolve is unlearnable. Occlusion is the obvious case: the gripper approaches, the object disappears behind it, and the policy is now acting on memory and hope. Less obvious is spatial grounding. Language-conditioned instructions like “put the mug behind the plate” require scene geometry a close-cropped wrist view never encodes.

The RoboMimic study (arXiv:2108.03298) quantified the general point years ago: on identical demonstrations, changing the observation space, including camera views, materially changed imitation learning outcomes. Observation design is a first-order variable. Yet teams routinely treat it as fixed plumbing while sweeping learning rates for a month.

We see the ceiling directly in our ablations at DexSet. On tabletop pick-place and insertion tasks, moving from wrist-only to wrist plus one calibrated external view produced the largest single jump in success rate we have measured from any data intervention. Adding a second external view helped again on occlusion-heavy tasks. No optimizer change came close.

What the Big Datasets Already Decided

The major post-2023 collection efforts treat multi-view as a requirement, and that consensus is evidence in itself. DROID (arXiv:2403.12945) recorded roughly 76,000 Franka episodes across 564 scenes with two external ZED stereo cameras and a wrist ZED Mini on every single episode, extrinsics calibrated. The protocol did not permit single-view shortcuts, because the authors understood that scene diversity is worthless if the model cannot see the scene.

Ego-Exo4D (arXiv:2311.18259) went further for human demonstration data: over 1,200 hours of skilled activity captured simultaneously from Aria glasses and multiple stationary exocentric cameras, synchronized and calibrated, precisely so models can learn correspondences between first-person and third-person views. Anyone planning to pretrain robot policies on human video needs that pairing.

Open X-Embodiment (arXiv:2310.08864) is the counterexample that proves the rule. It aggregates over a million trajectories from 22 embodiments, and its camera configurations are heterogeneous: wrist-only here, one exo view there, different poses everywhere, extrinsics often missing. Teams pretraining VLAs on it spend real engineering time coping with viewpoint inconsistency. The lesson is not that aggregation is bad; it is that viewpoint standards are cheap at capture time and expensive to retrofit.

The Real Bottleneck Is Operational, Not Scientific

The reason most data is still single-view is not ignorance; it is that multi-view capture is an operations problem disguised as a shopping list. Buying three cameras takes an afternoon. Keeping their extrinsics valid, their clocks aligned, and their output QA’d across hundreds of sessions is the part that defeats teams.

Three costs dominate:

  • Calibration maintenance. Extrinsics drift when mounts get bumped, booms sag, or thermal cycles shift fixtures. Without per-session verification, drift silently corrupts weeks of data. Our gate is a 0.5 px reprojection error check on a ChArUco sweep at every session start.
  • Synchronization. Software timestamps drift across devices; PTP (IEEE 1588) or hardware trigger lines fix it, but only if someone engineers and monitors the sync path. We reject sessions with more than 10 ms cross-camera skew on manipulation work.
  • Storage and throughput. A 4-camera 1080p30 rig produces 0.8 to 1.5 TB per capture day in our pipelines. Multiply your single-view storage budget by three or four, then add QA review time.

None of this is research. All of it is why the bottleneck persists.

What Closing the Gap Costs

The honest comparison is single-view capture cost versus multi-view capture cost versus the cost of the training runs and engineering time the single-view ceiling wastes. Our benchmark numbers:

Item Single-View (Wrist or 1 Exo) Multi-View (2 Exo + Wrist) Delta
Rig build $1,500 to $2,500 $4,500 to $7,000 +$3,000 to $4,500 one-time
Operated capture $18 to $25 / hr $26 to $38 / hr +$8 to $13 / hr
Storage per capture day ~0.3 TB ~1.0 TB ~3x
Occlusion-heavy task ceiling Hard cap, architecture-independent Removed The point

For a 500-hour dataset, the multi-view premium lands around $4,000 to $6,500 in capture plus the one-time rig delta. One senior engineer spending six weeks fighting a data-imposed plateau costs more, and one full retraining run on data you have to recollect anyway costs far more. The full cost model, camera comparisons, and rig geometry options are in our pillar guide: The Complete Guide to Exocentric & Multi-View Data for Robot Learning.

How to Scale Multi-View Capture Without Drowning

Scaling multi-view data means industrializing the boring parts. The checklist we run internally:

  • Standardize one rig geometry (we default to DROID-style: two external stereo, one wrist) so calibration procedures and QA gates are identical across stations.
  • Gate every session on a two-minute calibration verification clip; reject on reprojection error > 0.5 px or sync skew > 10 ms.
  • Automate extrinsics logging into the episode metadata, so every frame carries its camera poses forever.
  • Budget storage at 3-4x single-view and decide codec and retention policy before capture starts, not after the first full disk.
  • Ablate camera count on your own tasks before scaling past three views; in our experience the fourth camera rarely earns its cost.

Run the Failure Analysis Before the Next Sweep

If your policy metrics have plateaued and your dataset is single-view, run the failure analysis before the next architecture sweep: tag failures by whether the target was visible at decision time. If occlusion dominates, the fix is capture. Book a demo and we will walk you through calibrated multi-view sample episodes from our production rigs, with the calibration and sync metadata included.

Frequently Asked Questions

Why is multi-view data considered the bottleneck in physical AI?

Because model architectures and compute have outpaced data quality: policies trained on single-view episodes hit occlusion and spatial-grounding ceilings that no architecture change removes, and calibrated multi-view capture is operationally hard enough that most existing datasets never provided it.

In DexSet benchmarks, a DROID-style rig costs $4,500 to $7,000 versus $1,500 to $2,500 for single-view, and operated capture runs $26 to $38 per hour versus $18 to $25. Storage roughly triples.

Standardize one rig geometry across stations, gate every session on calibration and sync checks, embed extrinsics in episode metadata, budget storage at 3-4x single-view, and ablate camera count on your own tasks before adding a fourth view.

DROID enforced two external stereo views plus wrist across 76,000 episodes; Ego-Exo4D paired ego and exo video across 1,200+ hours for cross-view learning; RoboMimic showed observation space choices materially change imitation outcomes; Open X-Embodiment shows the integration cost when viewpoint standards are absent.

No. In our ablations the second view delivers the largest gain, a third helps on occlusion-heavy tasks, and a fourth is rarely distinguishable from noise while adding roughly 25 percent to storage and QA cost.

The Robotics Data Buyer’s Playbook: RFPs, Scorecards, and Pilots That Actually Work (2026 Guide)

QA_FAIL episode_0847.hdf5: sync_offset_max 41ms (cam_high vs qpos). That log line, or one very like it, is how many robot data programs discover they have a procurement problem: the flag fires weeks after signature, the contract never defined a sync tolerance, and the vendor technically delivered exactly what was ordered. The team can tell you its GPU budget to the dollar. It cannot tell you which of three vendors quoting $30, $45, and $80 per teleoperation hour would have caught that flag before shipping.

That gap exists for a structural reason. Robot training data has no established procurement discipline. Software buyers inherited decades of RFP practice from enterprise IT. Robotics data buyers inherited nothing, because the category barely existed before large-scale imitation learning and VLA models made demonstration data a line item worth millions. So teams improvise: they buy on price, on a demo video, or on whoever answered the email fastest. Then they discover, three months in, that 30 percent of delivered hours fail their own QA, the cameras were never time-synced to spec, and the contract says nothing about redelivery.

Our thesis, argued with numbers throughout this guide: buying robot data is buying a vendor’s pipeline, and because the defects that matter are latent, only three instruments filter them before signature: a written acceptance spec, a weighted scorecard, and a paid pilot. This guide gives you that procurement discipline, which the category is missing. You will get a 10-criterion weighted vendor scorecard, a 15-question RFP bank grouped by category, a red flags list built from real deals, and a 50-hour paid pilot protocol with pass/fail metrics. Everything here is designed to be copied into your next vendor evaluation.

We run teleoperation and egocentric capture pipelines at DexSet, which means we sit on the receiving end of RFPs every week. We have seen which questions separate serious buyers from tourists, and we have watched deals go sideways when those questions were never asked. This playbook is the document we wish every buyer sent us.

TL;DR: Key Takeaways

  • A robotics data buyer’s playbook is a structured process (RFP, weighted scorecard, red flags, paid pilot) for selecting and managing robot training data vendors.
  • Score vendors on 10 weighted criteria. Data quality SLAs, modality coverage, and calibration/sync spec carry the most weight (39 of 100 points combined).
  • Never sign an annual contract without a paid pilot. Our recommended protocol: 50 hours, measure usable-hour yield (target 85 percent or higher) and policy success delta on fixed eval tasks.
  • Pricing opacity is a signal, not an inconvenience. Vendors who publish per-hour ranges tend to survive QA audits; vendors who quote only after “discovery calls” often cannot.
  • Require delivery in a standard format (LeRobot, RLDS, or HDF5 with documented schema). A proprietary-format-only vendor is a lock-in risk.

What Is a Robotics Data Buyer’s Playbook?

A robotics data buyer’s playbook is a repeatable procurement process for evaluating, piloting, and contracting robot training data vendors, built around four artifacts: a request for proposal (RFP), a weighted vendor scorecard, a red flags checklist, and a pilot evaluation protocol. The playbook exists because robot demonstration data is a specification-heavy purchase, closer to contract manufacturing than to SaaS: what you receive is physical work (teleoperation, human video capture, annotation) frozen into files, and defects are expensive to detect after the fact.

The four artifacts map to four decisions:

  • RFP: which vendors are worth talking to at all.
  • Scorecard: how to compare the ones who respond, on the same axes, with weights that reflect your actual risk.
  • Red flags: which vendors to eliminate regardless of score.
  • Pilot protocol: whether the winning vendor’s production output matches their sample, before you commit annual budget.

Entity chain, stated plainly: your VLA model (OpenVLA, π0, GR00T-class) is trained by imitation learning on demonstration episodes; those episodes come from teleoperation rigs (ALOHA-style bimanual arms, VR-based systems) or egocentric human capture; the vendor’s calibration, sync, and QA pipeline determines whether those episodes are usable; your procurement process determines which vendor pipeline you inherit. Buying data is buying a pipeline.

Why Robot Training Data Is a Different Kind of Purchase

Robot training data procurement differs from every other data purchase because quality is defined by physics, not labels. In text or image annotation, a bad label is visible on screen. In robot data, the defects that kill model performance are invisible in a preview: a 40 ms sync offset between camera frames and joint states, extrinsics that drifted after someone bumped a wrist camera, gripper actions clipped at the edge of the calibration range. The episode looks fine. The policy trained on it does not.

Three properties follow from this:

Defects are latent. You often cannot see them until you train. This is why the pilot protocol below trains an actual policy instead of just eyeballing playback.

Specs are multidimensional. A single “hour of data” bundles camera count and placement (egocentric, exocentric, wrist), mono versus stereo, resolution and fps, depth, proprioception rate, action space, task diversity, and annotation depth. Two vendors quoting the same price are almost never quoting the same product.

Formats determine integration cost. The ecosystem has converged on a handful of formats: the LeRobot dataset format from Hugging Face (github.com/huggingface/lerobot), RLDS/TFDS episodes used by Open X-Embodiment (github.com/google-research/rlds; arxiv.org/abs/2310.08864), raw HDF5 in ALOHA conventions (arxiv.org/abs/2304.13705), and log formats like MCAP (github.com/foxglove/mcap) and rosbag2 (github.com/ros2/rosbag2) for ROS 2 stacks (docs.ros.org). A vendor who cannot deliver in at least one of these adds weeks of conversion work to every delivery.

Here is how the common delivery formats compare from a buyer’s perspective:

Format Origin / ecosystem Best for Random access Tooling maturity Buyer risk if this is the only option
LeRobot v2.x Hugging Face LeRobot Training loops, HF hub distribution Good (parquet + mp4) High, active community Low
RLDS / TFDS Google, Open X-Embodiment TF/JAX pipelines, OXE mixing Good High in TF world, thinner in PyTorch Low to medium
HDF5 (ALOHA-style) ACT / ALOHA papers Bimanual teleop episodes Good Medium, schema varies by lab Medium (demand a schema doc)
MCAP Foxglove Multimodal logging, ROS 2 Good Growing fast Medium
rosbag2 ROS 2 Robot-native logging Moderate High in ROS, low outside Medium
Proprietary format Single vendor Nothing you control Unknown None outside vendor High: lock-in, conversion cost, audit difficulty

The 10-Criterion Weighted Vendor Scorecard

The DexSet vendor scorecard evaluates a robotics data provider on ten criteria, each scored 1 to 5 and multiplied by a weight, for a maximum of 500 weighted points. The weights below reflect where we see deals actually fail; adjust them to your program, but keep quality, modality, and calibration at the top.

# Criterion Weight What a 5 looks like What a 1 looks like
1 Data quality SLAs 15 Contractual usable-hour yield (e.g., 90 percent pass your acceptance spec), free redelivery of failures “We do QA” with no numbers
2 Modality coverage 12 Ego + exo + wrist cams, stereo option, depth, synced proprioception, tactile roadmap Single exocentric mono camera
3 Calibration & sync spec 12 Published intrinsics/extrinsics per rig, time sync under 10 ms across streams, drift checks per shift “Cameras are calibrated” with no numbers or recheck cadence
4 Annotation accuracy 10 Documented rubric, dual-pass or audit sampling, 97 percent or higher on independent audit Unversioned labels, no audit trail
5 Throughput (hours/week) 10 Stated capacity per rig type, evidence of sustained delivery at your volume “We can scale to anything” without rig counts
6 Pricing transparency 10 Published per-hour ranges by rig and task complexity Pricing only after discovery calls
7 Privacy & consent chain 8 Signed operator/participant consent on file, face/bystander handling policy, GDPR posture No documented consent process
8 Format compatibility 8 Native LeRobot and RLDS export, HDF5/MCAP on request, schema docs Proprietary format only
9 Pilot terms 8 Offers paid pilot with buyer-defined acceptance criteria Refuses pilots or offers only cherry-picked samples
10 IP & licensing 7 You own delivered data and models trained on it; no reuse without consent Vendor retains broad reuse rights, ambiguous model ownership

How to use it: have two people score independently from the RFP responses and sample data, then reconcile. Anything below 350/500 exits the process. Anything scoring 1 or 2 on criteria 1, 3, or 10 exits regardless of total, because quality SLAs, calibration, and licensing are the three areas where a weak answer becomes an unrecoverable problem after signature.

The RFP Question Bank: 15 Questions That Do the Work

A robotics data RFP is a short, specific document (five pages beats fifty) whose questions force vendors to commit to numbers. Below are 15 questions we recommend, grouped by category. Vendors who answer all 15 with specifics belong on your shortlist. Vendors who answer with adjectives do not.

Quality and QA 1. What usable-hour yield do you commit to contractually against an acceptance spec we define, and what happens to hours that fail (redelivery, credit, or refund)? 2. Describe your per-episode QA pipeline: what is checked automatically, what is checked by humans, and what percentage of episodes get a second-pass review? 3. Provide annotation accuracy from your most recent independent audit, and the rubric it was measured against.

Hardware, calibration, and sync 4. For each rig type you would use on our program, list cameras (placement, mono/stereo, resolution, fps), depth sensors, and proprioception rates. 5. What is your maximum time-sync error across camera, joint-state, and action streams, and how is it measured and rechecked during production? 6. How often are intrinsics and extrinsics recalibrated, and do delivered episodes include per-rig calibration files?

Operations and throughput 7. How many rigs and operators would be dedicated to our program, and what sustained hours per week does that support at our spec? 8. What was your actual delivered volume for your largest program in the last six months (hours, not episodes)? 9. What is your process when task success rates drop or instructions change mid-program?

Commercial 10. Provide your per-hour pricing range by rig type and task complexity, and state every cost that is not included in it (setup, annotation, redelivery, format conversion). 11. What are your paid pilot terms: minimum hours, price, timeline, and whether buyer-defined acceptance criteria apply? 12. In which formats do you deliver natively (LeRobot, RLDS, HDF5, MCAP, rosbag2), and can you share a schema document today?

Legal and privacy 13. Who owns the delivered data and any models trained on it, and do you retain any right to reuse our episodes for other customers or your own models? 14. Describe your consent chain: what operators and any captured bystanders sign, and how consent records map to delivered episodes. 15. What happens to our task definitions, environment setups, and prompts after the engagement ends?

Red Flags: When to Walk Away

A red flag in robotics data procurement is any vendor behavior that predicts unrecoverable problems after contract signature. These are the ones we treat as disqualifying, based on deals we have watched from both sides:

  • Pricing available only after a discovery call. Opacity here usually means price is set by your budget, not their costs.
  • Refusal of a paid pilot with buyer-defined acceptance criteria. A vendor confident in their pipeline will take your money to prove it.
  • Sample data that is not from a production rig. Ask directly. Showcase rigs and production rigs can be different machines run by different people.
  • No stated time-sync tolerance. If they have never measured it, your training team will be the first to.
  • Proprietary delivery format with no export path. Lock-in plus audit difficulty in one package.
  • No consent documentation for operators or captured humans. This becomes your legal problem, not theirs.
  • Broad data reuse rights buried in the MSA. Your task distribution is competitive information.
  • “Unlimited scale” claims without rig counts. Throughput is rigs times shifts times yield. Anyone who will not show the multiplication is guessing.
  • No redelivery or credit policy for failed hours. QA without consequences is marketing.

Cost and Economics: What Robot Data Actually Costs

Robot training data pricing in 2026 clusters into ranges that depend on rig type, task complexity, and annotation depth, and any playbook needs those ranges to sanity-check quotes. The figures below are DexSet benchmarks and typical ranges we see across the market; treat them as calibration, not quotes.

Data type Typical market range (per hour) Main cost drivers Typical usable-hour yield we see
Bimanual teleoperation (ALOHA-class rig) $28–$60 Operator skill, task resets, rig count 80–90 percent
Humanoid on-robot teleoperation $150–$400 Robot cost and uptime, safety oversight, pilot skill 70–85 percent
Egocentric human video (mono to stereo + IMU) $15–$40 Participant recruiting, consent, headset hardware 80–90 percent
Exocentric multi-view capture $20–$50 Camera count, calibration, studio setup 80–90 percent
Dense annotation add-on +$8–$25 Label depth, audit sampling n/a
Independent QA add-on +$5–$12 Sampling rate, audit depth n/a

Two economics rules matter more than the sticker price:

Rule 1: price per usable hour, not per delivered hour. A $35/hr vendor at 70 percent yield costs you $50 per usable hour. A $44/hr vendor at 90 percent yield costs $48.89, arrives with less rework, and does not poison your training mix with borderline episodes.

Rule 2: model the pipeline, not the purchase. A 1,000-hour teleop program at $40/hr is $40,000 in data, but plan another 15 to 25 percent for integration, conversion, storage, and your own audit time if the vendor scores poorly on format compatibility and QA. Vendors who score 4 to 5 on criteria 1, 4, and 8 compress that overhead, which is usually worth more than a $5/hr discount.

The 50-Hour Paid Pilot Protocol

A pilot evaluation protocol is a fixed, paid, pre-contract engagement that measures whether a vendor’s production pipeline meets your acceptance spec, using metrics you define before the first hour is captured. We recommend 50 hours: large enough to expose process problems, small enough that a failed pilot costs weeks rather than quarters.

Protocol:

  • Fix the spec first. Write the acceptance spec (camera config, sync tolerance, task list, annotation rubric, delivery format) before contacting vendors. The pilot tests the vendor against the spec, not the spec against the vendor.
  • Pay for it. 50 hours at market rates is $1,500 to $4,500 for most modalities. Paying keeps the vendor’s incentives honest and gets you production treatment, not showcase treatment.
  • Demand production conditions. Same rigs, same operators, same QA pipeline that would run your annual contract. Put this in the pilot agreement.
  • Measure usable-hour yield. Run every delivered episode through your acceptance checks. Target: 85 percent or higher for teleop, 90 percent or higher for egocentric video.
  • Audit annotations on a sample. Independently re-label 5 percent of episodes. Target: 97 percent agreement or higher.
  • Train a policy and measure the delta. Fine-tune a fixed baseline (an ACT or diffusion policy head, or a small VLA) on your existing data alone, then on existing data plus pilot data. Evaluate both on the same fixed task set. The pilot passes only if success rate improves; flat or negative delta at 50 hours predicts flat or negative at 1,000.
  • Test the loader. Delivered data should load into your LeRobot or RLDS pipeline within one engineer-day. Log every schema surprise; each one recurs at scale.
  • Decide on numbers. Yield, audit accuracy, policy delta, loader time. Four numbers, agreed in advance, written into the pilot agreement.

Case Study: A Humanoid Foundation Model Team Rebuilds Its Buying Process

One humanoid foundation model team we work with came to us after a year of ad hoc purchasing: three vendors, three formats, no shared acceptance spec, and a training team spending roughly a day per delivery writing conversion scripts and triaging bad episodes. Their internal estimate was that a quarter of purchased hours never reached a training run.

They rebuilt procurement around the artifacts in this guide. The spec came first: stereo egocentric plus wrist cameras, sync under 10 ms, LeRobot delivery, a 60-task manipulation list. The RFP went to six vendors; four answered with numbers, two answered with adjectives and were dropped. The scorecard separated the four finalists cleanly, mostly on quality SLAs and licensing terms, where two vendors wanted broad reuse rights. Both finalists ran 50-hour paid pilots. One delivered 90 percent usable-hour yield and a measurable success-rate improvement on the fixed eval set; the other delivered 78 percent yield and failed the loader test.

The team signed with the first vendor. The numbers that mattered afterward: usable-hour yield across the first 1,000 contracted hours held within two points of the pilot, and data engineering time per delivery dropped from about a day to under an hour because the format and schema were locked before signature. No named clients, no invented revenue figures, just the pattern we see repeatedly: the pilot predicted production, because the pilot was run under production conditions.

Get the Template

The fastest way to apply this playbook is to not rebuild it. We packaged the full RFP question bank (the 15 questions above plus 20 more), the weighted scorecard as a spreadsheet with the math built in, the red flags checklist, and the pilot agreement language into a single download. Send the RFP as-is or strip it to the sections that match your program.

Related reading on this site: why procurement is the real bottleneck in physical AI, comparing data sourcing approaches and their costs, a humanoid team’s VLA data procurement case study, and five hidden challenges in robotics data RFPs.

Put the Playbook to Work

Ready to scope your data program? Talk to our team.

Frequently Asked Questions

What is a robotics data buyer’s playbook?

A robotics data buyer’s playbook is a structured procurement process for robot training data, built around four artifacts: an RFP with specification-forcing questions, a weighted vendor scorecard, a red flags checklist, and a paid pilot protocol with predefined pass/fail metrics.

Five to eight. Fewer gives you no basis for comparison; more than eight means your spec is probably too vague to have filtered anyone out, and evaluation time grows linearly with responses.

In our benchmarks, mature teleoperation pipelines deliver 80 to 90 percent of hours passing a well-defined acceptance spec. Below 80 percent, rework and training-mix contamination usually erase any price advantage.

Paid. A paid 50-hour pilot ($1,500 to $4,500 at typical market rates) obligates the vendor to run production processes and lets you enforce buyer-defined acceptance criteria. Free samples are curated by definition.

Require native delivery in at least one of LeRobot, RLDS/TFDS, or documented HDF5, with MCAP or rosbag2 as options for ROS 2 stacks. Reject proprietary-only formats; they create lock-in and make independent audits harder.

Start from the DexSet weights (quality SLAs 15, modality coverage 12, calibration/sync 12) and shift weight toward whichever criterion caused your last data problem. Keep quality, calibration, and licensing as automatic-disqualification criteria regardless of weights.

Download the RFP Template + Vendor Scorecard (XLSX) and send it to your shortlist this week, or book a demo and we will walk you through how DexSet answers all 15 questions, numbers included.

Robot Training Data Costs and Pricing: The Complete 2026 Guide

In our first year we priced a 6,000-hour teleoperation program the way most of this market still prices: by the raw hour. We hit our quoted rate and still broke the client’s budget model, because 22 percent of captured episodes failed their acceptance spec and nobody’s plan had funded the recollection. We rebuilt our cost model around that miss, and this guide is the rebuilt model, published.

The mistake was possible because robot training data has no commodity unit yet. An “hour” of data can mean a raw teleop stream with 30 percent unusable episodes, or a QA-passed, annotated, deduplicated hour that trains a policy. Those two hours differ in cost by 2x or more, and vendors quietly quote whichever one makes their number look better.

So here is the thesis this guide argues from the first table to the last: the only honest unit for pricing robot training data is cost per usable hour, and every quote you receive should be converted into that unit before you compare anything. To make the conversion possible, you will get our first-hand cost-per-hour benchmarks for teleoperation, egocentric human video, and multi-view exocentric capture; rig capex figures anchored to public hardware like ALOHA and UMI; the QA rejection math that separates raw hours from usable hours; and a worked budget for a 10,000-hour VLA data program.

DexSet collects egocentric, exocentric, teleoperation, mono, and stereo data for physical AI teams. We run these rigs, staff these operators, and eat these QA rejections every week. Every number below is either our own operating benchmark or a cited public source.

TL;DR: Robot Training Data Costs at a Glance

  • Teleoperation data: $28 to $60 per raw hour (rig amortization + operator + QA), based on DexSet benchmarks.
  • Egocentric human video: $15 to $40 per hour, the cheapest scalable modality.
  • Multi-view exocentric capture: $20 to $50 per hour depending on camera count and calibration load.
  • Annotation passes: $8 to $25 per hour extra, on top of any capture modality.
  • Rig capex: ~$20k for an ALOHA-style bimanual station (per the ALOHA paper), ~$32k for Mobile ALOHA, under $1k per UMI-style handheld gripper by our build estimates.
  • QA rejection runs 10 to 30 percent in our pipelines, so always budget on cost per usable hour, not raw hour.
  • Public scale references: Open X-Embodiment aggregates 1M+ trajectories across 22 embodiments; DROID contains 76k episodes.

What Do Robot Training Data Costs Actually Include?

Robot training data cost is the fully loaded price of producing one hour of demonstration data that a robot learning pipeline can actually consume, covering hardware amortization, operator or collector labor, QA review, annotation, and delivery infrastructure. Most published debates skip half of these line items, which is why budgets built from a single “per hour” quote fall apart in month two.

A defensible cost model has five layers:

  • Capture hardware (capex). Teleop stations, headsets, camera arrays, grippers. Amortized over 12 to 24 months of use.
  • Collection labor (opex). Teleoperators, camera-wearing collectors, scene setup staff. Usually the largest single line.
  • QA and curation (opex). Human review, automated checks, rejection and recollection. The most commonly omitted layer.
  • Annotation (opex). Language instructions, segmentation, keyframes, success labels. Priced per pass.
  • Infrastructure (opex). Storage, upload bandwidth, format conversion, dataset versioning.

When a vendor quotes you a single number, ask which of these five layers it covers. In our experience the quoted number usually covers layers 1 and 2 and quietly excludes 3 through 5, which add 30 to 80 percent on top.

Core Modalities and What Each One Costs

A data modality is the combination of sensor viewpoint and control method used to produce demonstrations: teleoperation, egocentric human video, exocentric multi-view capture, and their mono or stereo variants. Each modality has a distinct cost structure because each one shifts spend between hardware, labor, and QA differently.

Teleoperation Data: $28 to $60 per Hour

Teleoperation data is demonstration data produced by a human directly controlling a robot, typically through a leader-follower arm setup or a VR interface, so the recorded actions are executable robot trajectories. It is the gold standard for imitation learning and VLA post-training because actions come out in the robot’s own action space, but it is also the most expensive modality per hour.

Our benchmarks across bimanual manipulation programs:

  • Rig amortization: $4 to $9 per hour (a $20k to $32k station amortized over 18 months of two-shift use, including maintenance and spare grippers).
  • Operator labor: $18 to $38 per hour depending on region, task dexterity, and whether the task needs trained specialists (cable routing and garment handling sit at the top of that range).
  • QA overhead: $6 to $12 per hour, covering episode review, rejection, and partial recollection.

Total: $28 to $60 per raw teleop hour. Long-horizon mobile manipulation lands at the top of the range; tabletop pick-and-place with experienced operators lands at the bottom.

The entity chain matters here: teleoperation feeds imitation learning methods like ACT, which the ALOHA project introduced, and imitation learning at scale is what current VLA models are built on. The ALOHA paper (Zhao et al., “Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware,” https://arxiv.org/abs/2304.13705) demonstrated that a roughly $20k bimanual rig could produce data good enough for fine manipulation, which reset industry assumptions about capture hardware pricing.

Egocentric Human Video: $15 to $40 per Hour

Egocentric data is first-person video captured from head-mounted or body-mounted cameras while a human performs tasks with their own hands, giving models human-level dexterity examples without any robot in the loop. It is the cheapest modality to scale because the “rig” is a wearable and the collector is doing a task they already know how to do.

Our cost structure:

  • Hardware amortization: $1 to $4 per hour. Headsets and head-mounted cameras cost $300 to $3,500 and survive thousands of capture hours.
  • Collector labor: $10 to $26 per hour.
  • QA overhead: $5 to $10 per hour. Egocentric QA is dominated by motion blur, gaze drift, and occlusion checks.

Total: $15 to $40 per hour. The catch is the embodiment gap: human hands are not robot grippers, so egocentric data usually pretrains representations rather than directly supervising actions. Datasets like Ego4D and EgoExo4D (https://arxiv.org/abs/2311.18259) established the research value of this modality; the commercial question is purely about cost-effective volume.

Multi-View Exocentric Capture: $20 to $50 per Hour

Exocentric data is third-person video captured from fixed or mobile external cameras observing a task from multiple calibrated viewpoints, which gives models scene context and cross-view consistency that a single egocentric stream cannot. Cost scales with camera count and, more painfully, with calibration and synchronization labor.

Our benchmarks: $20 to $50 per hour for 3 to 8 synchronized views, including calibration checks at every scene change. Stereo pairs add roughly 15 to 25 percent over mono at the same view count because of the extra calibration and QA burden, but they buy you metric depth, which matters for manipulation policies.

Annotation: $8 to $25 per Hour, on Top of Everything Above

Annotation cost is the incremental price of adding structured labels to captured data: language instructions, subtask segmentation, success and failure flags, object masks, or keyframe tags. It is always additive to capture cost, and it is where “cheap” datasets quietly become expensive.

Typical per-pass pricing from our pipelines:

  • Language instruction labeling: $8 to $12 per data hour
  • Subtask segmentation: $10 to $16 per data hour
  • Dense object masks or contact annotation: $18 to $25 per data hour

Comparison Table: Cost per Hour by Modality

Modality Raw cost/hour (DexSet benchmark) Hardware amortization share QA rejection rate Best suited for
Teleoperation (bimanual) $28 to $60 $4 to $9 10 to 25% VLA post-training, imitation learning
Egocentric human video $15 to $40 $1 to $4 15 to 30% Representation pretraining, hand priors
Exocentric multi-view (3 to 8 cams) $20 to $50 $3 to $7 10 to 20% Scene understanding, cross-view learning
Stereo add-on (vs mono) +15 to 25% +$1 to $2 +2 to 5 pts Depth-dependent manipulation
Annotation pass (language) +$8 to $12 n/a n/a Instruction-following VLAs
Annotation pass (dense masks) +$18 to $25 n/a n/a Grasp and contact modeling

Rig Economics: Capex Benchmarks You Can Verify

Rig capex is the upfront hardware cost of a capture station before a single hour of data exists, and it is the number that determines whether building in-house ever beats buying data. The public record here is unusually good, so you do not have to trust vendor hand-waving.

Rig Approx. capex Source Notes
ALOHA bimanual teleop station ~$20,000 ALOHA paper, https://arxiv.org/abs/2304.13705 Leader-follower arms, cameras, frame
Mobile ALOHA ~$32,000 Mobile ALOHA paper, https://arxiv.org/abs/2401.02117 Adds mobile base for whole-body tasks
DexSet production teleop cell $20,000 to $32,000 First-hand DexSet build costs ALOHA-class arms plus industrial cameras, lighting, sync hardware
UMI handheld gripper Under $1,000 per unit (our build estimate) UMI paper, https://arxiv.org/abs/2402.10329 Portable gripper with wrist camera; no robot needed at capture time
Egocentric headset kit $300 to $3,500 Consumer/enterprise hardware pricing Camera glasses to mixed-reality headsets

Two practical lessons from running these rigs:

First, capex is rarely the problem. A $26k teleop cell running two shifts amortizes to under $9 per hour within 18 months. Labor and QA dominate every mature program we run.

Second, UMI-style handheld grippers changed the low end of the market. Because the capture device is a portable gripper with a wrist camera rather than a full robot cell, collection can happen in real homes and kitchens at egocentric-like labor costs while still producing gripper-centric trajectories. The trade-off is a heavier post-processing and QA load to recover clean actions.

Cost per Usable Hour: The Number That Actually Matters

Cost per usable hour is the total program spend divided by the hours that survive quality assurance, and it is always higher than the quoted cost per raw hour. This is the single most important correction to apply to any vendor quote, including ours.

In DexSet pipelines, QA rejection runs 10 to 30 percent depending on modality and task difficulty. Episodes get rejected for dropped frames, desynchronized views, failed task completion, occluded end-effectors, or annotation mismatches. The math is unforgiving:

Usable-hour math. At $40 per raw teleop hour with a 25 percent rejection rate, your real cost is $40 / 0.75 = $53.33 per usable hour. A competitor quoting $36 per hour with an unmeasured 35 percent rejection rate is actually charging $55.38. The cheaper quote is the more expensive dataset.

Quoted raw $/hr Rejection rate True cost per usable hour
$30 10% $33.33
$30 30% $42.86
$40 15% $47.06
$40 25% $53.33
$55 10% $61.11

When you evaluate any provider, require three things in writing: the measured rejection rate on a comparable program, who pays for recollection of rejected episodes, and whether QA review labor is inside or outside the quoted rate. If a vendor cannot produce a rejection rate, they are not measuring quality.

Budgeting a Program: A Worked Example

A data budget is a forward plan that converts a target usable-hour count into total spend across capture, QA, annotation, and infrastructure. Here is a realistic model for a VLA team that needs 10,000 usable teleop hours with language annotation.

  • Target: 10,000 usable hours
  • Assumed rejection rate: 20 percent, so raw capture target = 12,500 hours
  • Blended teleop rate: $42 per raw hour = $525,000 capture
  • Language annotation at $10 per usable hour = $100,000
  • Storage, versioning, delivery at roughly 4 percent of capture = $21,000
  • Total: ~$646,000, or $64.60 per usable annotated hour

For scale context, Open X-Embodiment pooled more than 1 million trajectories across 22 robot embodiments from 21 institutions (https://arxiv.org/abs/2310.08864), and DROID contains 76,000 episodes collected across 52 buildings (https://arxiv.org/abs/2403.12945). Those datasets exist because no single lab could afford to collect that volume alone, which tells you what the market already knows: collection cost, not model architecture, is the binding constraint on physical AI progress.

Build vs Buy: When Each One Wins

The build-vs-buy decision compares the fully loaded cost of standing up your own capture operation against a vendor’s cost per usable hour at your required volume and quality bar. Neither answer is always right; the crossover depends on volume, duration, and how much operational pain you can absorb.

Build wins when you need under roughly 2,000 hours of highly proprietary, robot-specific data, you already own the robots, and engineering time is genuinely available. Buy wins when you need volume and velocity: a vendor already amortized the rigs, trained the operators past the learning curve (operator throughput improves 30 to 50 percent over their first 200 hours in our programs), and built the QA tooling you would otherwise write from scratch. Most funded teams land on a hybrid: build one internal cell for rapid task iteration, buy production volume.

Case Study Proof: A Humanoid Foundation Model Team

A humanoid foundation model team came to us with a $400k data budget, a quoted competitor rate of $35 per hour, and a plan for 11,400 hours. The quote excluded QA review and carried no measured rejection rate. On a 200-hour pilot we measured 28 percent rejection against their own spec, which repriced the competitor dataset at $48.60 per usable hour before annotation.

We restructured the program: tightened the task spec to cut ambiguity-driven rejections, moved 30 percent of volume to egocentric capture for representation pretraining, and reserved teleop for post-training data. Result: 9,800 usable hours delivered inside the original budget, with rejection stabilized at 12 percent by week six. The lesson is not that our rate was lower. It is that cost per usable hour, measured on a pilot, is the only number that predicted their final spend.

Why Most Vendors Hide Pricing, and Why We Publish It

Hidden pricing is a deliberate market structure in which vendors quote deal by deal to maximize price discrimination, and it survives because buyers lack a shared benchmark. Large annotation-era incumbents built their margins on this asymmetry, and robot data inherited the habit.

We publish our ranges because the buyers we want, Heads of Data who run pilots and measure rejection rates, are exactly the buyers opaque pricing repels. Transparent ranges cost us the occasional overpriced deal and win us every buyer who has been burned before. You should treat any vendor’s refusal to publish even a range as information about how they expect the negotiation to go.

Free Download: The Robot Data Cost Model and RFP Scorecard

We packaged the math in this guide into two working documents: a cost model spreadsheet with editable assumptions for rejection rate, shift count, and amortization period, and a 24-question RFP scorecard covering the five cost layers, QA measurement, and recollection liability. Both are free, no email gate on the scorecard.

Related Reading

Put the Numbers to Work

Ready to scope your data program? Talk to our team.

Frequently Asked Questions

How much does robot training data cost per hour?

Based on DexSet’s operating benchmarks: teleoperation runs $28 to $60 per raw hour, egocentric human video $15 to $40, and multi-view exocentric capture $20 to $50. Annotation adds $8 to $25 per hour per pass. Divide any quoted rate by (1 minus the rejection rate) to get the true cost per usable hour.

The original ALOHA paper reports a bimanual rig built for roughly $20,000, and Mobile ALOHA extends it to whole-body mobile manipulation at roughly $32,000. Our production cells, with industrial cameras, lighting, and sync hardware added, land between $20,000 and $32,000.

In our pipelines, 10 to 30 percent of raw episodes fail QA, depending on modality and task complexity. Well-specified tabletop teleop can hold near 10 percent; long-horizon mobile tasks and loosely specified egocentric capture push toward 30 percent.

Building tends to win below roughly 2,000 hours of proprietary, robot-specific data when you already own robots and engineering time. Buying wins at volume because vendors have amortized rigs, trained operators, and existing QA tooling. Most teams run a hybrid.

Egocentric capture uses wearable cameras and human hands, so hardware costs hundreds to a few thousand dollars and collectors perform familiar tasks at natural speed. Teleoperation requires a $20k to $32k rig plus a trained operator, and outputs executable robot actions, which is what you pay the premium for.

Open X-Embodiment aggregates more than 1 million trajectories across 22 robot embodiments, and DROID contains 76,000 episodes. Both are useful pretraining anchors, but most teams still need proprietary data matched to their own embodiment and tasks.

Require the measured QA rejection rate on a comparable program, clarity on who pays for recollection, an itemized list of which cost layers the rate includes (hardware, labor, QA, annotation, infrastructure), and a paid pilot with your acceptance spec before any volume commitment.

Download the DexSet Robot Data Cost Model and RFP Scorecard, or book a 30-minute pricing walkthrough with our data operations team. We will run your task list through the same model we use internally and hand you the spreadsheet.

The Complete Guide to Exocentric & Multi-View Data for Robot Learning (2026)

You are three weeks from a milestone demo and the policy still drops the mug the moment the gripper crosses in front of it. The wrist camera looks perfect in the replay viewer, the demonstrations are clean, and none of it helps. The model never saw the scene from anywhere else, so the instant its one viewpoint goes blind, so does the policy.

The problem exists because most teams start with the camera that is easiest to mount, not the camera set that answers the questions their model will ask at inference time. A single egocentric or wrist view gives you fine-grained contact detail and nothing else: no scene context, no occlusion recovery, no spatial grounding for language-conditioned tasks. Fixing that after you have collected 400 hours of single-view teleoperation is expensive. Fixing it before you start costs a tripod and a calibration session.

The thesis of this guide is simple: camera geometry is a first-order training decision, not plumbing, and calibrated multi-view capture is the cheapest performance intervention most manipulation teams have not yet made. We will argue it three ways: with the published record (DROID, Ego-Exo4D, RoboMimic), with our own capture benchmarks and ablations, and with a worked customer example. This guide covers what exocentric and multi-view data actually is, which camera and rig configurations the major research datasets use, how to calibrate and synchronize multiple cameras without corrupting your dataset, and what all of it costs. The numbers come from our own capture operations at DexSet, where we run multi-camera teleoperation and human demonstration rigs daily for VLA and humanoid foundation model customers, plus the public dataset cards and papers we cite throughout.

TL;DR – Exocentric data is third-person footage of a robot or human performing a task; multi-view data captures the same episode from two or more calibrated, time-synchronized viewpoints at once. – The strongest public evidence for multi-view capture comes from DROID (two external stereo cameras plus a wrist camera on a Franka arm) and Ego-Exo4D (paired egocentric and exocentric video of skilled human activity). – Our benchmarks: a production-grade multi-view rig costs $2,500 to $12,000 to build, and calibrated multi-view capture runs $20 to $50 per hour depending on camera count, sync requirements, and QA depth. – The second camera delivers most of the benefit. In our ablations, going from one external view to two produced the largest jump in policy success; the fourth camera added little beyond storage cost. – Calibration drift and time sync are where multi-view datasets quietly die. Budget for recalibration checks every capture session, not every capture month.

What Is Exocentric Data for Robot Learning?

Exocentric data is visual data recorded from a third-person viewpoint that observes the robot or human demonstrator and the workspace from outside the body performing the task. A camera on a tripod behind the workbench, an overhead camera above a bin-picking cell, and a shoulder-height camera watching a humanoid fold towels are all exocentric views. The defining property is that the camera pose is independent of the actor’s motion.

Egocentric data is the complement: footage from the actor’s own perspective, such as smart glasses on a human demonstrator or a head-mounted camera on a humanoid. Wrist cameras sit in between; they move with the arm but do not share the actor’s gaze. Most serious manipulation stacks end up wanting both. The exocentric view supplies global scene context and object relationships, and the egocentric or wrist view supplies contact-level detail during grasps.

For a model, the practical difference shows up in failure modes. Policies trained only on exocentric views struggle with precise insertion because the interesting pixels are small and far away. Policies trained only on wrist views fail whenever the gripper blocks the object or the task requires reasoning about anything outside a 30 cm bubble. This is why datasets built for general-purpose policies pair the two rather than choosing.

What Is Multi-View Data?

Multi-view data captures a single episode from two or more cameras with known relative poses and aligned timestamps. Two properties separate a true multi-view dataset from a pile of videos that happen to point at the same table: calibrated extrinsics (each camera’s position and orientation relative to a shared frame, typically the robot base) and synchronization (frames across cameras correspond to the same instant, ideally within a few milliseconds).

Both properties are load-bearing. Without extrinsics, you cannot fuse views geometrically, project actions between frames, or train models that reason about 3D structure. Without sync, a policy learns from image pairs that show slightly different world states, which injects label noise you can never remove afterward. We reject capture sessions at DexSet when cross-camera timestamp skew exceeds 10 ms on manipulation tasks, because we have watched that skew turn into unexplainable policy failures downstream.

Multi-view is not the same as stereo. A stereo camera such as the ZED 2i is one viewpoint with two lenses on a fixed 12 cm baseline, built to estimate depth. A multi-view rig places whole cameras (mono, stereo, or RGB-D) at meaningfully different poses around the workspace. You can and often should combine them: two stereo cameras at different poses give you multi-view and per-view depth at once, which is exactly the configuration DROID chose.

What the Research Record Shows

The strongest public evidence for multi-view capture comes from three sources, and they agree with each other more than most robotics literature does.

DROID (arXiv:2403.12945, dataset) collected roughly 76,000 teleoperated episodes, about 350 hours, on Franka arms across 13 institutions and 564 distinct scenes. Every episode records two external ZED stereo cameras plus a wrist-mounted ZED Mini, with calibrated extrinsics. The authors made multi-view a hard requirement of the collection protocol, not an optional extra, because scene diversity only pays off if the model can see the scene.

Ego-Exo4D (arXiv:2311.18259, project page) is the largest paired ego-exo resource: over 1,200 hours of video of skilled human activity, captured simultaneously from Aria glasses on the participant and four or more stationary exocentric cameras, with camera calibration and time sync across the full rig. It exists because ego-only and exo-only datasets kept failing at cross-view tasks like translating “what I see” into “what the coach sees.” Robot learning inherits the same problem when transferring human video priors onto robot viewpoints.

Open X-Embodiment (arXiv:2310.08864, Hugging Face) aggregates over one million trajectories from 22 robot embodiments, and its per-dataset cards read like a natural experiment in camera configuration: some sources are wrist-only, some single-exo, some multi-view. Teams consuming it for VLA pretraining consistently report the friction of heterogeneous viewpoints, which is itself an argument for capturing calibrated multi-view from day one rather than harmonizing after the fact.

On the modeling side, the RoboMimic study (arXiv:2108.03298) found that observation space design, including which camera views feed the policy, materially changes imitation learning outcomes on the same demonstrations. Camera choice is a first-order hyperparameter, not plumbing.

Camera Hardware: What to Put on the Tripod

The camera decision is a trade between depth quality, shutter behavior, sync options, and price. These are the three units we deploy most, with the specs that actually matter for robot data capture.

Spec Intel RealSense D455 Stereolabs ZED 2i Luxonis OAK-D
Type Active IR stereo RGB-D Passive stereo + neural depth Stereo + RGB, on-device compute
Depth baseline 95 mm 120 mm 75 mm
Max RGB mode 1280x800 @ 30 fps 2208x1242 @ 15 fps (1080p @ 30) 4K RGB @ 30 fps (12 MP sensor)
Depth range (practical) 0.6 to 6 m 0.5 to 20 m (neural) 0.7 to 8 m
Shutter Global (depth), rolling RGB on some SKUs Rolling Global (OV9282 stereo pair)
IMU Yes Yes (9-DoF, barometer) Yes (on most variants)
Sync options External sync pin, multi-cam sync Timestamp-based, no genlock Hardware trigger via GPIO
Street price ~$420 ~$500 ~$250
Where it wins Tabletop manipulation, tight spaces Longer range, outdoor, mobile robots Budget rigs, embedded preprocessing

Two practical notes from our rigs. First, rolling shutter plus fast arm motion produces skewed geometry that quietly degrades any 3D supervision; if your task involves dynamic motion, weight global shutter heavily. Second, the D455 sync pin and the OAK-D hardware trigger let you drive frame capture from a shared signal, while the ZED relies on timestamp alignment, which is fine at 30 fps tabletop speeds and marginal for high-speed tasks.

Rig Geometry: Where the Cameras Go

Rig geometry is the arrangement of camera poses around the workspace, and it matters as much as the cameras themselves. The configurations below cover nearly everything we build.

Two external + wrist (the DROID pattern). One camera at roughly 45 degrees over each shoulder of the workspace, 0.8 to 1.2 m from the task center, plus a wrist camera. This is our default for tabletop manipulation. It gives occlusion recovery (when one external view is blocked, the other usually is not), stereo-of-stereos geometry for 3D checks, and contact detail from the wrist.

Overhead + front. An overhead camera looking straight down disambiguates object layout for pick-and-place and bin tasks; a front camera at torso height captures approach trajectories. Overhead mounts need rigid fixturing. A camera on a boom arm that sags 2 mm over a week of sessions will silently invalidate your extrinsics, which is one of the failure modes we cover in 5 Hidden Challenges in Exocentric & Multi-View Data.

Ego + exo paired (the Ego-Exo4D pattern). Glasses or head-mounted camera on the demonstrator plus stationary exocentric cameras. This is the configuration to choose when your training strategy includes human video pretraining, because it gives you the cross-view correspondence needed to transfer human priors to robot viewpoints.

Ring or arc arrays (4 to 8 cameras). Necessary for full-scene reconstruction, humanoid whole-body capture, or world-model training data. Expensive in sync engineering and storage, and in our experience unnecessary for most single-arm manipulation targets.

Calibration and Synchronization: The Unglamorous Core

Calibration is the process of estimating each camera’s intrinsics (lens and sensor parameters) and the rig’s extrinsics (relative camera poses, plus camera-to-robot-base transforms). Synchronization is making frame timestamps agree across cameras. Neither is hard the first time. Keeping both true across hundreds of capture hours is the actual job.

The standard toolchain is mature and free. OpenCV handles single and stereo calibration with checkerboard or ChArUco targets. Kalibr handles multi-camera and camera-IMU calibration, which you want the moment your rig has more than two cameras or a moving ego camera. ROS 2 ships camera_calibration in image_pipeline for in-place intrinsic calibration on live topics.

For sync, there are three tiers. Software timestamping (each camera stamps frames against system clocks) is free and drifts. PTP, the IEEE 1588 Precision Time Protocol, disciplines device clocks over Ethernet to sub-millisecond agreement and is the right answer for GigE machine vision cameras. Hardware genlock or trigger lines drive every sensor’s shutter from one signal and are the only way to get true same-instant exposure, which matters at high frame rates or with fast motion. Consumer RGB-D cameras mostly give you the first tier plus, on some units, a trigger pin; plan your rig around that constraint rather than discovering it later.

One more transform matters for robot data specifically: camera-to-robot-base, often called hand-eye calibration. Extrinsics between cameras tell you how views relate to each other; the camera-to-base transform tells you how all of them relate to the robot’s action space, which is what lets you project end-effector trajectories into any view or supervise 3D policies in the base frame. We estimate it by touching the arm’s end effector to known target points visible to the cameras, then verify against commanded poses. Skip this and your multi-view dataset is geometrically consistent with itself and disconnected from the robot.

Our operational rule: capture a 20-second calibration verification clip (a ChArUco board swept through the shared view volume) at the start of every session, and gate ingestion on reprojection error staying under 0.5 px and cross-camera skew under 10 ms. It costs two minutes per session and has saved entire capture weeks.

Cost and Economics: What Multi-View Actually Costs

Multi-view capture costs break into rig build (one-time) and operated capture (per hour). The figures below are our own benchmarks from rigs we run at DexSet; treat them as typical ranges, not quotes.

Configuration Rig Build Cost Operated Capture Cost Typical Use
1 external RGB-D + wrist (baseline) $2,500 to $4,000 $20 to $28 / hr Prototyping, single-task policies
2 external stereo + wrist (DROID-style) $4,500 to $7,000 $26 to $38 / hr VLA training data, tabletop manipulation
Ego + 3 exo paired (Ego-Exo4D-style) $6,000 to $9,500 $32 to $44 / hr Human demo capture, cross-view learning
6-camera arc, hardware-triggered $9,000 to $12,000 $40 to $50 / hr Humanoid whole-body, reconstruction, world models

Rig build includes cameras, mounts and fixturing, sync hardware, a capture workstation, and the first calibration. Operated capture includes the operator, session calibration checks, QA review, and annotation-ready packaging; it excludes task design and annotation itself. Storage is the line item teams forget: a 4-camera rig at 1080p30 generates roughly 0.8 to 1.5 TB per capture day depending on codec, and multi-view triples or quadruples whatever single-view budget you had.

The marginal-value curve is the key economic fact. In our ablations on tabletop pick-place and insertion tasks, adding the second external view to a wrist-only setup produced the largest single improvement in policy success. The third camera helped mainly on occlusion-heavy tasks. The fourth was rarely distinguishable from noise while adding about 25 percent to storage and QA cost. Buy the second camera before you buy anything else; justify the fourth with an experiment, not a hunch. We break the per-configuration trade-offs down further in Comparing Exocentric & Multi-View Approaches: Pros, Cons & Costs.

Case Study: Multi-View at Scale for a VLA Team

A humanoid foundation model team came to us with a policy that plateaued at roughly 60 percent success on cluttered tabletop tasks, trained on single-view teleoperation data. Their failure analysis pointed at occlusion: success dropped sharply whenever the target object was blocked from the lone camera during the approach.

We rebuilt their capture around a DROID-style rig: two external stereo cameras plus wrist, hardware-checked calibration each session, sub-10 ms sync, and a QA gate on reprojection error. Over eight weeks we delivered several hundred hours of calibrated multi-view episodes across their task list. Retrained on the multi-view data with the same architecture and episode count, the policy’s success on the occlusion-heavy split improved by double digits, and their engineers stopped hand-labeling “camera blocked” failure cases entirely. The full pipeline detail is in How We Scaled Multi-View Data for a VLA Model, and the strategic version of the argument is in Why Multi-View Data Is the Biggest Bottleneck in Physical AI.

Downloadable: The Multi-View Rig RFP Scorecard

If you are evaluating data vendors or scoping an internal rig, the questions that separate real multi-view capability from marketing are specific: How are extrinsics verified per session? What is your cross-camera timestamp skew tolerance and how is it measured? What reprojection error gates ingestion? What is the storage format and per-hour deliverable size? We have packaged these into a one-page RFP scorecard with scoring weights you can hand to procurement. Download it below; it is the same rubric we hold our own rigs to.

Put This Guide to Work

If your current dataset is single-view and your policy failures cluster around occlusion or spatial grounding, the fix is a capture decision, not an architecture search. Download the Multi-View Rig RFP Scorecard to evaluate vendors or your own rig plan, or book a demo to see calibrated multi-view episodes from our production rigs, including sample data you can load the same day.

Frequently Asked Questions

What is exocentric data in robot learning?

Exocentric data is third-person visual data captured by cameras positioned outside the robot or demonstrator, observing the actor and workspace from fixed external viewpoints. It provides global scene context that egocentric and wrist cameras cannot, and it pairs with those views in most modern manipulation datasets.

Two calibrated external views plus a wrist camera cover most manipulation use cases; this is the configuration DROID used across 76,000 episodes. In our ablations the second external camera delivers the largest gain, while a fourth camera rarely improves policy success enough to justify its storage and QA cost.

Stereo is one viewpoint with two lenses on a fixed short baseline, designed to estimate depth. Multi-view places separate cameras at meaningfully different poses around the workspace with calibrated extrinsics. A rig can be both, for example two ZED 2i stereo cameras at different positions.

Based on our benchmarks at DexSet, multi-view rig builds run $2,500 to $12,000 depending on camera count and sync hardware, and operated calibrated capture runs $20 to $50 per hour including session calibration checks and QA. Storage adds roughly 0.8 to 1.5 TB per capture day for a 4-camera 1080p30 rig.

Three tiers: software timestamping (free, drifts), PTP / IEEE 1588 clock discipline over Ethernet (sub-millisecond, right for GigE machine vision cameras), and hardware genlock or trigger lines (true same-instant exposure). For tabletop manipulation at 30 fps, keep cross-camera skew under 10 ms and verify it every session.

Ego-Exo4D is the largest, with over 1,200 hours of simultaneously captured ego (Aria glasses) and exo (stationary camera) video of skilled human activity, with calibration and sync across the rig. In robot-collected data, DROID pairs a wrist (near-ego) view with two calibrated external stereo views.

VLA models benefit from multi-view data because language-conditioned tasks require scene-level grounding and manipulation requires contact-level detail, which no single viewpoint provides. The RoboMimic study showed observation space choices materially change imitation learning outcomes, and Open X-Embodiment’s heterogeneous camera setups are a recurring friction point for teams pretraining VLAs.

Comparing Exocentric & Multi-View Data Approaches for Robot Learning: Pros, Cons & Costs

A buyer on a scoping call last month put the field’s confusion into one sentence: everyone tells him to collect multi-view data, and nobody will tell him which multi-view. He was right to push. “Multi-view” describes at least five distinct capture strategies with different rig costs, different failure modes, and different value per training hour, and the honest answer, the thesis of this post, is that the right configuration is determined by your task list and training strategy, not by your budget or by whichever public dataset you read about first. Choosing the wrong one is not a small mistake. A team that builds a six-camera arc when their tasks needed a wrist camera and one tripod has burned rig budget, tripled their storage bill, and slowed capture throughput for nothing.

The confusion is understandable. Public datasets each embody one choice without explaining the alternatives: DROID picked two external stereo cameras plus wrist, Ego-Exo4D picked glasses plus stationary exo arrays, and Open X-Embodiment inherited whatever its 22 source labs happened to mount. The papers report what was captured, not the decision tree.

This post is that decision tree. We compare the five approaches we quote and build most often at DexSet, with honest pros, cons, and cost ranges from our own rigs, and end with a matrix mapping task types to configurations. For the underlying camera specs, calibration toolchain, and sync engineering, see the pillar guide: The Complete Guide to Exocentric & Multi-View Data for Robot Learning.

Key Takeaways – Five capture approaches dominate: single exo + wrist, DROID-style (2 exo + wrist), ego+exo paired human capture, dense arrays (4-8 cameras), and sim-rendered multi-view. – DROID-style is the default for tabletop manipulation and VLA training data: $4,500 to $7,000 rig, $26 to $38 per operated hour in our benchmarks. – Ego+exo paired capture is the only approach that supports human-video pretraining with cross-view transfer; it costs more in sync engineering than in cameras. – Sim-rendered views are nearly free per view but inherit the sim-to-real gap; they complement real capture, they do not replace it.

The Five Approaches, Defined

A capture approach is the combination of camera count, camera placement, actor type (robot or human), and sync method used to record training episodes. The five that cover almost every real program:

  • Single exocentric + wrist. One fixed external camera plus a wrist camera on the robot. The minimum viable multi-view setup.
  • DROID-style: two exocentric stereo + wrist. Two external stereo cameras (ZED 2i class) at distinct poses plus a wrist camera, calibrated extrinsics, as used across DROID’s 76,000 episodes (arXiv:2403.12945).
  • Ego + exo paired human capture. Glasses or head-mounted camera on a human demonstrator plus stationary exocentric cameras, the Ego-Exo4D pattern (arXiv:2311.18259).
  • Dense array (4-8 cameras). Hardware-triggered ring or arc around the workspace for reconstruction-grade coverage.
  • Sim-rendered multi-view. Arbitrary virtual cameras rendered from simulation, optionally mixed with real data.

Master Comparison Table

Approach Rig Build Capture Cost / Hr Sync Difficulty Occlusion Coverage Human-Video Pretraining Main Risk
Single exo + wrist $2,500 to $4,000 $20 to $28 Low Partial No Blind spots remain
DROID-style (2 exo + wrist) $4,500 to $7,000 $26 to $38 Moderate Good No Calibration upkeep
Ego + exo paired $6,000 to $9,500 $32 to $44 High (moving ego cam) Good Yes Ego-exo time alignment
Dense array (4-8 cams) $9,000 to $12,000 $40 to $50 High (trigger/genlock) Excellent No Storage, diminishing returns
Sim-rendered multi-view Compute only ~$1 to $5 equivalent None Perfect Limited Sim-to-real gap

Rig and capture figures are DexSet benchmarks, including session calibration checks and QA; sim figures are rough GPU-time equivalents.

Where Each Approach Wins and Loses

Single exo + wrist earns its place as a starting point. Pros: cheapest real multi-view, simple calibration (one extrinsic pair), enough to break the wrist-only occlusion ceiling for many tasks. Cons: one blocked view and you are back to single-view; no view redundancy for QA cross-checks. We recommend it for prototyping and single-task policies, and we recommend planning the mount points for camera two on day one.

DROID-style is the workhorse, and not by accident. Two external views mean occlusion of one is usually covered by the other; three total views give the RoboMimic-style observation flexibility that lets ML teams ablate view combinations later (arXiv:2108.03298). Cons: per-session calibration verification becomes mandatory, because three cameras drift three ways. In our operations the added QA overhead is roughly 5 percent of session time. This is what we quote when a VLA team asks for a default.

Ego + exo paired solves a different problem: it is the only configuration that produces the ego-exo correspondences needed to pretrain on human demonstration video and transfer to robot viewpoints, the exact gap Ego-Exo4D was built to close. Pros: human demonstrators are fast and cheap per episode; the data doubles as a bridge to large human-video corpora. Cons: the ego camera moves, so extrinsics to the world frame change every frame and must be recovered via SLAM or the glasses’ own tracking; time alignment between glasses and fixed cameras is the hardest sync problem on this list. Choose it when your training strategy explicitly includes human video.

Dense arrays buy near-complete coverage and reconstruction-grade geometry for humanoid whole-body work and world-model data. The cons compound quietly: hardware triggering or genlock is effectively mandatory, storage runs 3 to 4 TB per capture day at 1080p30 in our pipelines, and, in every ablation we have run on single-arm manipulation, cameras five through eight never moved the success metric. Buy this coverage for reconstruction, not for policy learning on tabletop tasks.

Sim-rendered multi-view costs almost nothing per additional view, which is genuinely useful for view-invariance augmentation and architecture prototyping. But every rendered view inherits the simulator’s gap in contact dynamics, materials, and lighting. Teams in the Open X-Embodiment consortium (arXiv:2310.08864) mix sim and real rather than substituting one for the other, and that matches our experience: sim views stretch a real multi-view dataset, they do not replace it.

Decision Matrix: Match the Approach to the Program

Your Situation Recommended Approach
Prototyping one task, tight budget Single exo + wrist, mounts pre-planned for a second exo
Training VLA / manipulation foundation data at scale DROID-style (2 exo + wrist)
Pretraining on human demonstrations or video Ego + exo paired
Humanoid whole-body, reconstruction, world models Dense array, hardware-triggered
Need view diversity beyond rig budget DROID-style real capture + sim-rendered augmentation

One category the table cannot capture: switching costs. Moving from single-exo to DROID-style mid-program is cheap if the mount points and calibration workflow were planned for it, and painful if they were not, because your existing episodes and your new episodes will differ in geometry and your training pipeline has to reconcile them. Moving from robot-only capture to ego+exo is a bigger jump; it changes your demonstrator pool, your sync architecture, and your annotation scheme at once. Teams that expect to make either move should write the target configuration into their schema now, even if the extra cameras arrive next quarter.

Two cross-cutting rules. First, whatever you choose, log extrinsics and sync metadata into every episode; the approach you pick today is the aggregation problem someone inherits in two years. Second, ablate before you scale: run 20 hours in the candidate configuration, train, and let the success metric pick the rig.

Turn the Matrix Into a Procurement Rubric

If you are scoping a capture program or comparing vendors, download the Multi-View Rig RFP Scorecard. It turns this decision matrix into weighted evaluation questions on calibration verification, sync tolerances, and deliverable formats, the same rubric we hold our own rigs to.

Frequently Asked Questions

What is the cheapest way to get multi-view robot data?

A single external camera plus a wrist camera, at roughly $2,500 to $4,000 for the rig and $20 to $28 per operated capture hour in DexSet benchmarks. It breaks the wrist-only occlusion ceiling for many tasks but leaves blind spots a second external view would cover.

For VLA training data and tabletop manipulation at scale, usually yes: the second external view covers occlusions the first misses and enables view ablations later. The premium over single-exo is about $2,000 to $3,000 in rig cost and $6 to $10 per hour.

No. Rendered views are nearly free and useful for view-invariance augmentation, but they inherit the simulator’s gaps in contact dynamics, materials, and lighting. Production programs mix sim views with real calibrated capture rather than substituting.

When your training plan includes learning from human demonstration video. Paired capture, as in Ego-Exo4D, provides the cross-view correspondences needed to transfer first-person human priors to third-person robot viewpoints.

Task-dependent, but in our single-arm manipulation ablations, cameras beyond the third stopped moving policy success while adding roughly 25 percent storage and QA cost per view. Dense arrays of 4 to 8 cameras are justified for reconstruction and whole-body humanoid work, not tabletop policies.

The Complete Guide to Egocentric Data Collection for Robotics (2026)

3,670 hours. That is the complete Ego4D corpus, the largest first-person video dataset ever assembled (arXiv:2110.07058), and it amounts to roughly seven months of one person’s waking life. The robot foundation models expected to generalize across every kitchen, warehouse, and workbench on earth are drawing from a first-person data supply about that size, while their language-model cousins trained on trillions of tokens. The shelf is not thin. It is nearly bare, and teleoperation refills it at a few hundred action-labeled hours per rig per year.

The gap exists because robots perceive the world from their own body. A vision-language-action (VLA) model driving a humanoid needs to learn from footage that looks like what its head camera will actually see: hands entering the frame from below, objects at counter height, occlusions caused by the manipulator itself. That viewpoint is called egocentric, and until recently there was no scaled, systematic way to collect it. Our thesis, argued with numbers throughout this guide: egocentric capture is the only collection method that scales to foundation-model volumes, but it earns that scaling only when modality mix, viewpoint geometry, and annotation depth are derived from the training mechanism rather than from a hardware catalog.

This guide covers the full stack: what egocentric data collection for robotics actually is, the modalities that matter (mono, stereo, depth, IMU, gaze, hand pose), the hardware options from a $500 Quest 3 to Aria Gen 2 research glasses, how the data plugs into policy training, and what it costs per hour. We include the benchmark numbers we use internally at DexSet, because pricing opacity is the single biggest complaint we hear from buyers.

DexSet supplies egocentric, exocentric, teleoperation, mono, and stereo data to physical AI teams. We have built and rebuilt the capture rigs, the QA pipelines, and the annotation stacks described below, and most of the numbers in this guide come from our own production logs.

TL;DR: Key Takeaways – Egocentric data collection captures first-person visual and sensor streams from a camera mounted at the head or chest of a human (or robot), matching the viewpoint a robot policy will see at inference time. – It is the most scalable source of manipulation pretraining data: a human wearing glasses collects demonstrations 3 to 5 times faster than a teleoperator on a bimanual rig, at roughly one third to one half the cost per hour in our benchmarks. – Hardware ranges from ~$500 (Meta Quest 3, GoPro head mounts) to research-grade Aria Gen 2 glasses with calibrated multi-camera, IMU, eye tracking, and on-device machine perception. – Egocentric human video does not replace teleoperation; the strongest results (EgoMimic, co-training pipelines behind modern VLAs) combine both. The action gap between human hands and robot grippers is the core technical problem. – Our production benchmarks: raw egocentric capture runs $15 to $22 per hour; fully annotated (hand pose, object tracks, temporal segmentation) runs $30 to $40 per hour. Teleoperation runs $28 to $60 per hour depending on rig and task complexity.

What Is Egocentric Data Collection for Robotics?

Egocentric data collection for robotics is the practice of recording synchronized video and sensor streams from a first-person viewpoint, typically a head-mounted or chest-mounted camera worn by a human demonstrator, to train robot perception and control models. The defining property is viewpoint: the camera sees the scene the way an embodied agent sees it, with the demonstrator’s own hands and workspace in frame.

Three properties separate egocentric robotics data from ordinary first-person video:

  1. Sensor completeness. A YouTube cooking clip is RGB only. A robotics-grade egocentric recording carries calibrated camera intrinsics and extrinsics, IMU streams for ego-motion, and often stereo pairs or depth so that 3D structure can be recovered.
  2. Action recoverability. The footage must support extraction of what the hands did: 3D hand pose, object 6-DoF tracks, contact events. Without recoverable actions, egocentric video is only useful for representation pretraining, not policy learning.
  3. Task intent. Recordings are organized into episodes with defined start states, goals, and outcomes, mirroring how robot demonstration datasets like those in Open X-Embodiment are structured (arXiv:2310.08864).

The reference datasets here are Meta’s Ego4D, 3,670 hours of daily-life egocentric video across 74 locations (arXiv:2110.07058), and EgoExo4D, which pairs egocentric and exocentric views of skilled activities with dense annotations (arXiv:2311.18259). Both were built for video understanding research; robotics teams now treat them as the template for what scaled first-person capture looks like.

Egocentric vs. Exocentric: Why Viewpoint Determines Value

Exocentric data is footage captured from an external, third-person viewpoint, such as a tripod camera watching a workbench, while egocentric data is captured from the agent’s own point of view. The distinction matters because a policy trained purely on third-person views must solve an extra correspondence problem at deployment: mapping an external observation of a scene onto its own body frame.

The relationship chain that matters for buyers runs like this: egocentric human video teaches visuomotor priors, teleoperation (through systems like ALOHA, arXiv:2304.13705) provides robot-embodiment action labels, imitation learning consumes both, and VLA models such as OpenVLA (arXiv:2406.09246) and π0 (arXiv:2410.24164) sit at the top of the stack. EgoExo4D demonstrated why you often want both viewpoints of the same episode: the exocentric view disambiguates whole-body motion that the egocentric camera cannot see.

In our pipelines, paired ego-exo capture adds roughly 20 to 30 percent to per-hour cost (a second calibrated camera, cross-view sync, extra QA) and is worth it for whole-body humanoid work. For tabletop manipulation, egocentric plus a single fixed reference camera is usually sufficient.

Core Modalities in Egocentric Capture

A modality is one synchronized sensor stream within a recording, and the modality mix determines both the cost of capture and what training objectives the data can support. The five that come up in nearly every RFP we see:

Mono RGB. A single color stream is the cheapest to capture and the only modality most internet-scale pretraining uses. Sufficient for representation learning and video prediction, insufficient on its own for metric 3D.

Stereo RGB. Two horizontally offset cameras allow metric depth recovery through disparity. Stereo is the workhorse for manipulation because grasp points need metric accuracy. Devices like the Intel RealSense D435i and D455 provide hardware-synced stereo pairs plus an onboard IMU; the D455’s wider baseline (95 mm vs. 50 mm) improves depth accuracy at counter-to-room distances.

Depth. Active or computed depth gives per-pixel range directly. Active IR depth degrades in sunlight and on reflective surfaces, which is why most of our outdoor captures rely on passive stereo instead.

IMU. Accelerometer and gyroscope streams recover head motion, enabling ego-motion compensation and SLAM. Project Aria glasses carry two IMUs precisely because ego-motion is that important for downstream 3D reconstruction (projectaria.com).

Gaze and hand pose. Eye tracking (available on Aria) reveals attention targets before the hand moves, and 3D hand pose is the raw material for retargeting human demonstrations to robot grippers. These are the modalities that convert “video” into “demonstration.”

Hardware: The 2026 Egocentric Rig Landscape

An egocentric capture rig is the wearable hardware package (cameras, IMU, compute, mounting) used to record first-person data, and rig choice is the largest single driver of both data quality and program cost. The four setups we run or evaluate most often:

Rig Approx. Hardware Cost Sensors Video Spec (Typical Capture Config) Calibration Best For
Meta Quest 3 (Passthrough Capture) ~$500 Stereo RGB passthrough, IMU, hand tracking 1280×1280 per eye class, 30 fps effective capture Factory, limited access Budget hand-tracked demos, teleop UI doubling as capture
Aria Gen 2 Research Glasses Research program device (not retail) RGB + mono SLAM cameras, 2 IMUs, eye tracking, spatial mics, on-device hand tracking RGB up to 8 MP class, SLAM cams at high frame rate Full factory calibration + MPS services Research-grade egocentric corpora, gaze + hand pose at scale
GoPro Head/Chest Mount $350–$550 Mono RGB (wide FOV), IMU Up to 5.3K, we typically run 4K/60 or 2.7K/120 Self-calibrated (checkerboard) High-volume, low-cost mono capture; harsh environments
RealSense D435i/D455 Helmet Rig (Custom) $700–$1,200 built Stereo IR + RGB, active depth, IMU 848×480 depth at 90 fps or 1280×720 at 30 fps, RGB 1080p Manual, per-rig Metric depth for manipulation, sim-to-real alignment

Three field notes from running these at volume:

  • Quest 3 is underrated as a capture device because its hand tracking gives you approximate 3D hand pose for free, but passthrough capture access is constrained and image quality trails dedicated cameras.
  • Aria Gen 2 is the quality ceiling. Factory-calibrated multi-camera plus eye tracking plus machine perception services means far less post-processing on our side. Access runs through Meta’s research program rather than retail channels, which affects fleet scaling plans.
  • GoPro rigs win on ruggedness and unit economics. The cost is downstream: no depth, so you pay in annotation and 3D lifting compute instead of hardware.

How Egocentric Data Trains Robot Policies

Egocentric data enters robot learning through three mechanisms: representation pretraining, action retargeting, and co-training with robot demonstrations. Understanding which mechanism you are buying data for should drive every spec decision.

Representation pretraining. Visual encoders pretrained on large egocentric corpora like Ego4D transfer to manipulation tasks better than encoders trained on third-person or object-centric images, because the visual statistics (hands, near-field objects, ego-motion blur) match deployment. This is the lowest-risk use of egocentric data: mono RGB is enough, and annotation requirements are light.

Action retargeting. Human hand trajectories extracted from egocentric video are mapped onto robot end-effectors, turning passive video into pseudo-demonstrations. This requires recoverable 3D hand pose, which is why gaze-and-hand-instrumented devices matter. EgoMimic (arXiv:2410.24221) showed that egocentric human data captured on Aria glasses, combined with robot data, improves manipulation policies over robot data alone.

Co-training. Modern VLA training mixes robot episodes (teleop, in formats like the LeRobot dataset standard, github.com/huggingface/lerobot) with human egocentric episodes in one curriculum. The human data supplies breadth of scenes and objects; the robot data anchors the action distribution to the target embodiment. Cross-embodiment training in Open X-Embodiment established the pattern that heterogeneous data mixtures beat single-source datasets, and egocentric human video is the cheapest heterogeneity you can add.

The failure mode to respect: the embodiment gap. Human wrists have degrees of freedom robot grippers lack, human reach and eye height differ from most robot platforms, and human demonstrators exploit compliance no rigid arm has. Data collection protocols can shrink this gap (constrained grasps, robot-plausible motion instructions, matched camera height), and we bake those constraints into our capture scripts.

Collection Approaches Compared: In-House, Crowdsourced, Vendor

A collection approach is the operational model used to produce the data: who wears the rig, who designs the tasks, and who owns QA. Most teams land on one of three models, and the trade-offs are stable across every program we have run.

Approach Cost per Finished Hour (Our Benchmarks) Throughput Ramp Quality Control Where It Breaks
In-house Capture Team Typically well above vendor rates once salaries, rig fleet, and management overhead are loaded in; often roughly double Slow: 2–3 months to steady state Tight, iterative Scaling past ~10 collectors; hiring drag
Crowdsourced / Distributed Low headline rate, before rejection Fast but noisy Weak; rejection rates at or beyond the top of our 10 to 30 percent planning band are common in our audits Calibration, sync, task compliance
Specialist Vendor (DexSet Model) $15–$40 fully QA’d, annotation-dependent 2–4 weeks to first delivery Contractual, sampled + automated Task designs needing daily iteration with your researchers

The honest read: in-house wins when your task distribution changes weekly and researchers need to redesign protocols on the fly. A vendor wins when the task list is stable and the bottleneck is volume with consistent QA. Crowdsourcing looks cheap until you price the rejection rate and the engineering time spent triaging unsynced, uncalibrated footage.

Cost and Economics: What Egocentric Data Actually Costs

The cost of egocentric data is best expressed as dollars per finished, QA-passed hour at a defined annotation depth, because raw capture is a minority of total program cost. Competitors rarely publish numbers, so here are ours. These are current DexSet benchmark ranges, stated as typical figures we see across programs, not quotes:

Line Item Typical Range (Per Finished Hour) Notes
Raw Egocentric Capture (mono/stereo, IMU, episode structure) $15–$22 Collector time, rig amortization, upload, storage
+ Temporal Annotation (task/step segmentation, outcome labels) +$5–$8 Largely tooling-assisted
+ 3D Hand Pose + Object Tracks +$8–$12 The expensive layer; drives the $30–$40 fully-annotated figure
Paired Ego + Exo Capture +20–30% on capture line Second camera, cross-view sync, extra QA
Teleoperation (for comparison) $28–$60 Rig and task complexity dependent; bimanual fine manipulation sits at the top

Two planning rules of thumb from our production logs:

  • Budget 15 to 25 percent of hours for QA failure. Motion blur, dropped IMU packets, and off-task episodes are facts of life. Vendors should absorb this; if you collect in-house, plan for it.
  • Annotation depth should follow the training mechanism. If you are pretraining encoders, do not pay for hand pose. If you are retargeting actions, hand pose is the whole point. We regularly see RFPs over-specified by $10+ per hour because annotation depth was copied from a paper rather than derived from the training plan.

One more line item buyers forget: storage and delivery. Stereo capture at 30 fps with IMU sidecars generates roughly 50 to 120 GB per hour depending on resolution and compression, so a 5,000-hour corpus is a few hundred terabytes before derivatives. Cloud egress on a corpus that size is real money, which is why we quote delivery format and transfer method inside the per-hour price rather than as a surprise on the final invoice. Ask any vendor to do the same.

At these rates, a 5,000-hour egocentric corpus with full annotation lands between $150K and $200K. The equivalent volume via bimanual teleoperation would run $140K to $300K and take 3 to 5 times as long on a comparable rig fleet, which is the arithmetic behind the current industry shift toward human egocentric pretraining with a smaller teleop fine-tuning set.

Case Study Proof: 4,000 Hours for a Humanoid VLA Team

A humanoid foundation model team came to us with a stalled co-training experiment: their teleop corpus was high quality but topped out near 400 hours, and scaling it 10x on their own rigs would have taken most of a year. We scoped a 4,000-hour egocentric program across kitchen, warehouse shelving, and assembly-bench task families, captured on stereo rigs at camera heights matched to their robot’s head frame, with hand pose and object tracks on the 30 percent of hours their researchers flagged as retarget-critical.

Delivery ran 14 weeks. Their team reported that co-training on the mixed corpus improved task success on unseen-object manipulation evaluations relative to their robot-only baseline, consistent with the direction published in EgoMimic-style co-training work. The full breakdown of task families, QA gates, and the capture protocol is in the case study blog this pillar links to below.

Related reading from this series: – Why Egocentric Data Collection for Robotics Is the Biggest Bottleneck in Physical AIComparing Egocentric Data Collection Approaches: Pros, Cons and CostsCase Study: Scaling Egocentric Data Collection for a VLA Model5 Hidden Challenges in Egocentric Data Collection

Download: The Egocentric Data RFP Template

An RFP template turns this guide into a procurement tool: it lists the 40 questions we believe every buyer should ask a data vendor, covering rig specs, calibration evidence, sync tolerances, QA sampling methodology, annotation rubrics, pricing structure, and data licensing. We built it from the RFPs we answer, including the questions we wish more buyers asked. Download it, delete our name from the header if you like, and send it to every vendor on your shortlist including us.

Put These Numbers to Work

If you are scoping an egocentric data program this quarter, two options. Download the RFP template and pressure-test every vendor with it, or book a 30-minute demo and we will walk you through sample episodes from our stereo and Aria-class rigs, including the QA reports we ship with every batch. Either way, you leave with real numbers instead of a sales deck.

Frequently Asked Questions

What is egocentric data collection for robotics?

Egocentric data collection for robotics is the recording of synchronized first-person video and sensor streams (RGB, stereo, depth, IMU, gaze, hand pose) from head- or chest-mounted rigs, structured into task episodes, to train robot perception and manipulation models.

Egocentric data captures a human performing tasks with their own hands from a first-person camera, while teleoperation data captures a robot performing tasks under human control, with exact joint-space action labels. Egocentric data is cheaper and faster to scale; teleoperation data matches the robot embodiment exactly. Most modern VLA pipelines use both.

Common rigs include Meta Quest 3 (~$500, stereo passthrough and hand tracking), Aria Gen 2 research glasses (calibrated multi-camera, IMU, eye tracking), GoPro head or chest mounts ($350 to $550, mono RGB), and custom helmet rigs built around Intel RealSense D435i or D455 stereo depth cameras.

In DexSet’s benchmarks, raw QA-passed egocentric capture runs $15 to $22 per hour, and fully annotated data with hand pose and object tracks runs $30 to $40 per hour. Teleoperation data runs $28 to $60 per hour for comparison.

No. Egocentric human video scales pretraining and improves generalization, but the embodiment gap between human hands and robot grippers means policies still need robot-embodiment data (teleoperation or autonomous rollouts) for reliable control. Research such as EgoMimic supports combining both.

It depends on the mechanism: encoder pretraining benefits from thousands of hours of lightly annotated video, while retargeting pipelines often start showing gains with hundreds of hours of densely annotated, task-matched capture combined with a robot demonstration set.