Skip to main content

Dexset

Case Study: How a Series B Humanoid Startup Rebuilt Data Procurement for Its VLA Model

Grasp success on the team’s fixed eval suite dropped six points in a single week, and for two days nobody could explain it. The model had not changed. The training code had not changed. What had changed, it eventually emerged, was a data batch: one of their three vendors had silently rescaled gripper actions between deliveries. The team had money, talent, and a working humanoid platform. What stalled their VLA program was procurement: three data vendors, three delivery formats, no shared acceptance spec, and an ML team that had quietly stopped trusting incoming data. By their own estimate, roughly a quarter of purchased hours never reached a training run.

Nothing about that situation was unusual. It is what happens when data purchasing grows by accretion: a vendor added for a demo push here, another for a task category there, each with its own definitions of “hour,” “annotated,” and “calibrated.” The result is not one data pipeline but several, with the training team as the manual integration layer.

This post walks through how the team rebuilt procurement in one quarter using the four artifacts from the Robotics Data Buyer’s Playbook: a spec-first RFP, a weighted scorecard, a red flags screen, and a 50-hour paid pilot. The client is anonymized (a Series B humanoid startup; we do not name customers), but the process and the numbers are the real pattern we see when buyers switch from improvisation to protocol. The thesis this case argues: pilots run under production conditions predict production performance, and paper evaluation alone does not.

DexSet was one of the vendors evaluated in this process, so we saw the RFP, the pilot, and the contract from the receiving side. That is exactly the vantage point a buyer should want documented.

Key Takeaways

  • Starting point: three vendors, three formats, ~74 percent average usable-hour yield, about one engineer-day of triage per delivery.
  • The rebuild: one written acceptance spec, an RFP to six vendors, a 10-criterion weighted scorecard, paid 50-hour pilots for the top two.
  • Pilot results diverged sharply: 90 percent yield with a positive policy delta versus 78 percent yield and a failed loader test.
  • Twelve weeks later: yield held within two points of the pilot across the first 1,000 contracted hours, and per-delivery engineering time dropped from a day to under an hour.

The Starting Point: Accretion, Not Architecture

The team’s data supply had accreted rather than been designed: each vendor relationship was created for a short-term need, and none was ever re-evaluated. Vendor A delivered undocumented HDF5 with a schema that shifted between batches. Vendor B delivered rosbag2 that had to be converted and re-timestamped. Vendor C delivered mp4s plus CSV joint logs with no sync guarantees at all.

The measurable symptoms, from their own tracking:

  • Average usable-hour yield around 74 percent across vendors, measured against the acceptance checks they eventually wrote down.
  • Roughly one engineer-day per delivery spent on conversion and triage.
  • Two training regressions in one quarter traced back to data batches, one to a sync drift issue, one to inconsistent gripper action scaling between vendors.

The trigger for change was mundane: a planning meeting where nobody could answer what an incremental 1,000 hours of manipulation data would actually cost, because nobody could say what fraction would be usable.

Step 1: Write the Spec Before Talking to Anyone

The acceptance spec is a short document that defines what a usable hour means for a specific program, and writing it was the most valuable week of the whole rebuild. Theirs fit on three pages: stereo egocentric plus wrist cameras at defined resolutions and frame rates, proprioception at 50 Hz minimum, time sync under 10 ms across all streams, a 60-task manipulation list with reset criteria, an annotation rubric, and delivery in LeRobot format with a schema doc, following the conventions the ecosystem inherited from Open X-Embodiment and the ALOHA line of work (arxiv.org/abs/2310.08864; arxiv.org/abs/2304.13705).

The immediate effect was that vendor quotes became comparable for the first time. The secondary effect was internal: the ML and data teams had to agree on what they were actually buying, which surfaced two silent disagreements about camera placement that had been corrupting cross-vendor consistency for months.

Step 2: RFP and Scorecard

The RFP went to six vendors with fifteen questions requiring numeric answers: committed usable-hour yield, sync tolerance and recheck cadence, audit accuracy, rig counts behind throughput claims, per-hour pricing by task complexity, pilot terms, and licensing. Four vendors answered with numbers. Two answered with adjectives and were dropped without further calls, which is the red flags screen doing its job cheaply.

The four remaining responses were scored independently by two people on the 10-criterion weighted scorecard (quality SLAs 15, modality coverage 12, calibration/sync 12, down to IP/licensing at 7). Independent scoring earned its keep in the reconciliation meeting: the two scorers disagreed by more than a point on only three criteria, and every disagreement traced to an ambiguous vendor answer rather than a difference in judgment. Those ambiguities went back to the vendors as written clarification requests, which is a politer and more useful outcome than one person’s optimism deciding the ranking.

The scoring separated the field more on contract terms than on hardware: two vendors wanted broad rights to reuse delivered episodes for other customers, which the team treated as disqualifying on the licensing criterion. Two finalists advanced.

Step 3: The 50-Hour Paid Pilots

Both finalists ran paid 50-hour pilots under production conditions, judged on four numbers fixed in the pilot agreements before capture began. This is the step that paper evaluation cannot replace, and the results diverged more than the scorecards had predicted.

Pilot metric (agreed in advance) Threshold Finalist 1 Finalist 2
Usable-hour yield vs acceptance spec ≥85% 90% 78%
Annotation accuracy (independent 5% re-label) ≥97% 98.2% 96.1%
Policy success delta on fixed 12-task eval Positive +6 points vs baseline +1 point
Loader time into LeRobot pipeline ≤1 engineer-day ~2 hours Failed (schema mismatches, 3+ days)

The policy-delta test deserves a note. The team fine-tuned the same fixed baseline policy twice, once on their existing mix, once with the 50 pilot hours added, and evaluated both on the same 12 tasks. Fifty hours is a small delta at VLA scale, and they treated it that way: not as proof of final performance, but as a canary. Data that helps at 50 hours might help at 1,000; data that does nothing at 50 hours is a warning you can act on before signing an annual contract.

Step 4: Contract and the Twelve Weeks After

The contract encoded the pilot numbers rather than replacing them with prose: yield commitment at 88 percent with redelivery of failures, sync tolerance and recalibration cadence as delivery requirements, LeRobot schema versioned in an appendix, quarterly re-audit rights, and no vendor reuse of delivered episodes.

The results over the first 1,000 contracted hours were unremarkable in the best sense. Yield held between 88 and 90 percent, within two points of the pilot. Engineering time per delivery fell from roughly a day to under an hour, because the format and schema were locked before signature instead of negotiated after each batch. The two silent disagreements about camera placement did not recur, because the spec, not tribal memory, was the reference.

The honest caveat: this process cost about five weeks end to end and roughly $7,000 in pilot fees across two finalists. For a team spending six figures annually on data, that insurance premium was around 5 percent of program cost. For a team buying 20 exploratory hours, it would be overkill, and we say so in the full playbook.

Next Step

Ready to scope your data program? Talk to our team.

Frequently Asked Questions

How much did the procurement rebuild cost compared to what it saved?

About five weeks of process time and roughly $7,000 in paid pilot fees. Against it: yield moved from ~74 to ~90 percent, and per-delivery engineering time dropped from a day to under an hour, which on a six-figure annual program repaid the process cost within the first deliveries.

Because scorecards rank paper, and paper can lie by omission. In this case the two finalists scored within a few points of each other, then diverged by 12 yield points and a failed loader test under production conditions.

It is enough to judge the vendor’s pipeline, not the model’s ceiling. Treat the policy success delta as a canary: positive movement at 50 hours justifies scaling; zero movement is a cheap warning.

The pilot numbers themselves: yield commitment with redelivery terms, sync tolerance as a delivery requirement, a versioned schema appendix, re-audit rights, and an explicit ban on vendor reuse of delivered episodes.

Four Ways to Buy Robot Training Data, Compared: Pros, Cons, and Costs

Taped above a monitor in a buyer’s office we visited last year was a coffee-stained, three-page acceptance spec, its 10 ms sync tolerance circled twice in red pen. That battered document was doing more procurement work than the forty-page RFP folder on the shelf beside it, because it forced every vendor quote onto the same axes. It also marked its owners as unusual: teams rarely frame “how are we going to buy this” as a decision at all. They email two vendors someone met at CoRL, pick the cheaper quote, and only discover they chose a procurement approach when it fails.

The failure is predictable because each buying approach has a known cost structure and a known blind spot. An informal purchase is fast and blind. A full RFP is thorough and slow. A pilot-first approach measures what matters but covers one vendor at a time. Open datasets are free and almost never match your embodiment or task distribution.

This post compares the four approaches on speed, cost, and risk, with the math that lets you pick deliberately. It draws on the same scorecard and pilot protocol as our full Robotics Data Buyer’s Playbook, which is where the reusable templates live.

At DexSet we respond to all four buying styles weekly, so we see their outcomes from the supplier side: which approaches produce clean contracts and which produce disputes about what “an hour of data” was supposed to mean.

Key Takeaways

  • There are four common procurement approaches: informal purchase, full RFP + scorecard, pilot-first, and open-data-plus-top-up. Each has a distinct cost and risk profile.
  • Informal buying is cheapest to run ($0 process cost) and most expensive to survive: yield surprises routinely add 20 to 40 percent to effective cost.
  • The RFP + scorecard + pilot combination costs roughly 3 to 5 weeks and $1,500 to $4,500 in pilot fees, and is the only approach that measures quality before annual commitment.
  • Open datasets (Open X-Embodiment, DROID, Ego4D) are excellent for pretraining and mixing, but embodiment and task mismatch means most teams still purchase targeted data on top.

Approach 1: Informal Purchase

An informal purchase is vendor selection without a written specification, scoring method, or pilot: the buyer requests quotes, reviews sample clips, and signs with the most convincing option. It is how most first data purchases happen, and it is defensible exactly once, at very small volume, when you are still learning what to specify.

Pros: fastest path to first data (days, not weeks); no process overhead; fine for exploratory volumes under 20 hours.

Cons: samples are curated, so latent defects (sync offsets, calibration drift) go undetected; quotes are not comparable because no shared spec exists; no contractual yield commitment, so failed hours are your loss; format surprises arrive with the first delivery.

Cost profile: zero process cost up front. In our benchmarks, yield surprises and conversion work typically add 20 to 40 percent to effective per-usable-hour cost versus a piloted vendor. At 1,000 hours, that is $8,000 to $16,000 of avoidable spend on a $40/hr program.

Approach 2: Full RFP with Weighted Scorecard

A full RFP approach sends a written acceptance spec and a fixed question set to five to eight vendors, then scores responses on weighted criteria before any commitment. This is classic procurement discipline adapted to robot data: quality SLAs, calibration and sync specs, throughput evidence, pricing transparency, licensing terms.

Pros: quotes become comparable because everyone bids the same spec; weak vendors self-eliminate (in our experience, roughly a third of recipients answer with adjectives instead of numbers); the scorecard creates an audit trail for the decision; licensing and consent problems surface before signature.

Cons: takes two to four weeks; still paper-based, so a vendor can score well and underdeliver; overkill below roughly $25,000 in annual data spend.

Cost profile: the process costs internal time only, typically 20 to 30 person-hours across spec writing, scoring, and reconciliation. It buys you comparability and eliminates the worst outcomes, but on its own it does not measure production quality.

Approach 3: Pilot-First

A pilot-first approach skips broad solicitation and goes straight to a 50-hour paid pilot with one or two candidate vendors, judged on predefined metrics: usable-hour yield, annotation audit accuracy, policy success delta on a fixed eval set, and loader time into LeRobot or RLDS. It optimizes for measured evidence over paper promises.

Pros: measures the only thing that matters, production output; small, fixed downside ($1,500 to $4,500 per pilot at market rates); fast when you already know the credible vendors; the policy-delta test catches defects no document review can.

Cons: covers only the vendors you pilot, so a better option may never be evaluated; sequential pilots take longer than parallel paper scoring; requires you to have a stable eval task set and baseline policy, which very early teams may lack.

Cost profile: $3,000 to $9,000 to pilot two vendors, plus about one engineer-week for evaluation. Expensive compared to reading PDFs, cheap compared to one bad quarter of deliveries.

Approach 4: Open Data Plus Targeted Top-Up

The open-data approach builds the base training mix from public corpora, then purchases only the targeted data the public sets cannot provide. The public layer is genuinely strong now: Open X-Embodiment spans over one million episodes across 22 embodiments (arxiv.org/abs/2310.08864), DROID adds 76,000 diverse teleop episodes (arxiv.org/abs/2403.12945), and Ego4D provides thousands of hours of egocentric human video (arxiv.org/abs/2110.07058), most of it accessible through Hugging Face dataset cards and RLDS tooling.

Pros: near-zero acquisition cost for pretraining scale; well-documented formats (RLDS, LeRobot conversions); community-validated quality.

Cons: embodiment mismatch (your gripper, camera placement, and control rates differ from the source robots); task distribution rarely matches your product; licenses vary and some restrict commercial use, so legal review is not optional; fine-tuning still demands in-domain demonstrations, which puts you back in one of the first three approaches for the data that moves your metrics most.

Cost profile: storage and engineering only for the public layer, then standard market rates ($28 to $60 per teleop hour, $15 to $40 per egocentric hour in our benchmarks) for the top-up volume, which is typically 10 to 30 percent of total hours but drives most of the task-specific performance.

Side-by-Side Comparison

The four approaches differ most in where they spend money: process time up front, or rework after delivery.

Approach Time to contract Process cost Quality measured before commitment? Typical effective cost penalty vs piloted baseline Best for
Informal purchase 3–10 days ~$0 No +20–40% Exploratory buys under 20 hours
Full RFP + scorecard 2–4 weeks 20–30 person-hours Partially (paper only) +5–15% Annual spend above $25k, multiple candidate vendors
Pilot-first 2–3 weeks per vendor $1,500–$4,500 per pilot Yes Baseline Teams with stable eval tasks and known vendor shortlist
Open data + top-up 1–2 weeks (legal + integration) Engineering time Yes for public layer, no for top-up unless piloted Depends on top-up approach Pretraining scale plus targeted fine-tuning

The pattern most mature buyers converge on is a hybrid: RFP to filter the field, scorecard to rank it, pilot to verify the winner, open data underneath it all as the pretraining base. That sequence is exactly what the Robotics Data Buyer’s Playbook packages, including the scorecard weights and pilot pass/fail thresholds.

One sequencing note from the supplier side: run the approaches in that order, not in parallel. Teams that pilot before writing a spec end up measuring vendors against criteria invented after the data arrived, which makes the results unarguable in exactly the wrong way; nobody can agree what a pass looks like. Teams that RFP without a spec get six incomparable quotes and mistake the spread for market variance. The spec is upstream of everything, takes about a week to write, and is the only artifact in the process that costs nothing but attention.

Next Step

Ready to scope your data program? Talk to our team.

Frequently Asked Questions

Which robot data procurement approach is cheapest overall?

The RFP-plus-pilot hybrid, once volume passes roughly $25,000 per year. Informal buying has the lowest process cost but the highest effective cost, because yield surprises add 20 to 40 percent on typical programs.

Not fully. Open X-Embodiment, DROID, and Ego4D are strong pretraining bases, but embodiment and task mismatch means fine-tuning still needs in-domain demonstrations, typically 10 to 30 percent of total hours purchased to spec.

For exploratory volumes under about 20 hours, where the goal is learning what to specify rather than feeding a production training run. Anything feeding a release model deserves at least a pilot.

Send the RFP to five to eight vendors, score all responses, then pilot the top one or two. Piloting more than two rarely changes the decision and doubles the evaluation load.

Four numbers agreed in advance: usable-hour yield (target 85 percent or higher), annotation accuracy on an independently re-labeled 5 percent sample (97 percent or higher), policy success delta on a fixed eval set, and loader time into your training format (one engineer-day or less).

Why Data Procurement, Not Data Capture, Is the Biggest Bottleneck in Physical AI

Two teams approached us in the same quarter with nearly identical VLA data programs. The first sent a two-page acceptance spec and an RFP that demanded numbers, then ran a paid pilot before signing with anyone. The second compared three price sheets and took the lowest. Half a year later, the first team’s deliveries were entering training runs the day they arrived; the second team was still bisecting a training regression that traced back to a sync tolerance no contract had ever specified. Same budget class, same architecture, divergent quarters. The difference was not capture quality. It was procurement.

The uncomfortable part is that capture itself has scaled. ALOHA-class rigs are reproducible from public documentation (arxiv.org/abs/2304.13705). Open X-Embodiment pooled over a million episodes across 22 embodiments (arxiv.org/abs/2310.08864). DROID collected 76,000 teleop episodes across 13 institutions (arxiv.org/abs/2403.12945). The hardware and process knowledge exist. What has not scaled is the buying side: most teams still purchase demonstration data with less rigor than they apply to a laptop refresh.

This post makes the case that procurement is now the binding constraint, shows what the gap costs in numbers, and gives you the four artifacts that close it. It condenses the full Robotics Data Buyer’s Playbook, which includes the complete scorecard and RFP question bank.

We see this from the vendor side at DexSet. The buyers who send a real spec get better data at better prices than the buyers who send a budget and a hope, because a real spec lets us commit to numbers instead of hedging against unknowns.

Key Takeaways

  • Capture capacity has commoditized; vendor selection has not. Weak procurement is now the most common cause of stalled VLA data programs.
  • The cost of a bad pick is measured in usable hours: a cheap vendor at 70 percent yield can cost more per usable hour than a pricier one at 90 percent, once rework is priced in.
  • Latent defects (sync error, calibration drift) are invisible in previews and only surface in training, which is why sample-based buying fails.
  • The fix is procedural, not heroic: a spec-first RFP, a 10-criterion weighted scorecard, a red flags list, and a 50-hour paid pilot.

What Makes Procurement the Bottleneck

The procurement bottleneck is the delay and waste created when robot data purchasing decisions are made without a specification, a scoring method, or a pilot, forcing quality problems to surface downstream during training. It shows up as three concrete failure patterns.

Pattern one: the invisible defect. Robot data defects are latent. A 40 ms sync offset between camera frames and joint states will not appear in video playback, but it corrupts the state-action mapping your imitation learning policy depends on. Teams that buy on sample previews systematically miss this class of problem, then spend weeks bisecting training regressions that were purchased, not coded.

Pattern two: the spec vacuum. When the buyer has no written acceptance spec, every vendor quote describes a different product. One vendor’s “hour of manipulation data” is stereo egocentric plus wrist cameras with dense annotations; another’s is a single 720p exocentric mono stream. Comparing their prices is meaningless, and procurement stalls in clarification loops that a two-page spec would have prevented.

Pattern three: the format tax. Deliveries arrive in whatever the vendor uses internally: undocumented HDF5, half-converted rosbag2, a proprietary container. Your engineers pay the conversion cost into LeRobot or RLDS on every delivery. In our experience that tax runs 15 to 25 percent of program cost when format compatibility was never contracted.

The Cost of Buying Badly, in Numbers

The cost of weak procurement is best expressed as price per usable hour, which is the delivered price divided by the fraction of hours that pass your acceptance spec. The sticker price is the number vendors compete on; the usable-hour price is the number your training run experiences.

Scenario Sticker price Usable-hour yield True price per usable hour 1,000-hour program cost (usable basis)
Cheapest bid, no pilot $32/hr 70% $45.71 $45,710
Mid bid, sample-only check $40/hr 82% $48.78 $48,780
Higher bid, passed 50-hr pilot $46/hr 90% $51.11 $51,110
Cheapest bid after rework and triage time $32/hr + engineering time 70% $55–$60 effective $55,000–$60,000

Read the last row carefully. The cheapest vendor is the most expensive one once you price the engineering time spent triaging failures and patching training mixes, and that is before counting the schedule slip, which no spreadsheet captures but every roadmap feels. These are typical ranges from our benchmarks; your yields will vary, which is exactly why you measure them in a pilot instead of assuming them.

The Fix: Four Artifacts, Not More Meetings

The fix for the procurement bottleneck is a set of four reusable artifacts that convert vendor selection from judgment calls into measurements. Each one is small. Together they remove the guesswork that creates the bottleneck.

  • A spec-first RFP. Write the acceptance spec (cameras, sync tolerance under 10 ms, task list, annotation rubric, delivery format) before contacting anyone. Then send 15 questions that force numeric answers: committed usable-hour yield, audit accuracy, rig counts, per-hour pricing by task complexity.
  • A weighted scorecard. Ten criteria, weights summing to 100, with data quality SLAs, modality coverage, and calibration/sync spec carrying 39 points between them. Two independent scorers, reconciled. Automatic disqualification for weak answers on quality, calibration, or licensing.
  • A red flags list. Some behaviors end the conversation regardless of score: pricing only after discovery calls, refusal of paid pilots, proprietary-only formats, no consent documentation, broad data reuse rights.
  • A 50-hour paid pilot. Production rigs, production operators, and four predefined numbers: usable-hour yield (85 percent or higher), annotation audit accuracy (97 percent or higher on a 5 percent sample), policy success delta on a fixed eval set, and loader time into your LeRobot or RLDS pipeline (one engineer-day or less).

The full versions, including the complete scorecard weights and all 15 RFP questions, are in the Robotics Data Buyer’s Playbook.

How to Tell If Procurement Is Your Bottleneck

A procurement bottleneck diagnosis is a check of where data-program time actually goes, and it takes one honest hour with your delivery logs. Run through five questions:

  • Can anyone state your acceptance spec from memory, or point to the document? If the spec lives in tribal knowledge, every vendor conversation is renegotiating it implicitly.
  • What fraction of delivered hours reached a training run last quarter? If nobody tracks this number, assume it is worse than you think; teams that start measuring usually find 15 to 30 percent of purchased hours never trained anything.
  • How long does a delivery take to enter the pipeline? More than a few hours of engineer time per batch means you are paying the format tax on every delivery.
  • Could you defend your current vendor choice to your board with numbers? A scorecard produces that defense as a byproduct. A demo-based decision cannot.
  • When did a data defect last cost you a training debugging cycle? If the answer is “this quarter,” the defect entered through procurement, not through capture.

Two or more uncomfortable answers means the bottleneck is upstream of your training code, and the fix below is cheaper than the symptom.

Scaling the Playbook Across Programs

Scaling this process means running the same artifacts on every purchase rather than reinventing evaluation per deal. The spec becomes a living document versioned alongside your model releases. The scorecard weights shift as your risks shift: teams early in data collection weight modality coverage higher; teams scaling a proven recipe weight throughput and SLAs higher. Pilot results accumulate into an internal vendor database, which is the closest thing this industry has to a track record. After three or four cycles, vendor evaluation drops from weeks of meetings to days of scoring, and that is the bottleneck removed.

Next Step

Ready to scope your data program? Talk to our team.

Frequently Asked Questions

Why is procurement a bigger bottleneck than capture in physical AI?

Because capture methods are now well documented (ALOHA, DROID, Open X-Embodiment) while buying practices are not. Latent defects like sync error survive sample review and surface in training, so unstructured purchasing converts directly into stalled training runs.

Delivered price divided by the fraction of hours passing your acceptance spec. A $32/hr vendor at 70 percent yield costs $45.71 per usable hour, which is why sticker-price comparisons mislead.

Around 50 paid hours. That volume exposes process problems (calibration drift, operator variance, QA gaps) while keeping a failed pilot cheap: typically $1,500 to $4,500 at market rates.

Native delivery in LeRobot, RLDS/TFDS, or documented HDF5, with MCAP or rosbag2 as ROS 2 options. Proprietary-only delivery is a disqualifier because it adds conversion cost to every delivery and blocks independent audits.

The Robotics Data Buyer’s Playbook: RFPs, Scorecards, and Pilots That Actually Work (2026 Guide)

QA_FAIL episode_0847.hdf5: sync_offset_max 41ms (cam_high vs qpos). That log line, or one very like it, is how many robot data programs discover they have a procurement problem: the flag fires weeks after signature, the contract never defined a sync tolerance, and the vendor technically delivered exactly what was ordered. The team can tell you its GPU budget to the dollar. It cannot tell you which of three vendors quoting $30, $45, and $80 per teleoperation hour would have caught that flag before shipping.

That gap exists for a structural reason. Robot training data has no established procurement discipline. Software buyers inherited decades of RFP practice from enterprise IT. Robotics data buyers inherited nothing, because the category barely existed before large-scale imitation learning and VLA models made demonstration data a line item worth millions. So teams improvise: they buy on price, on a demo video, or on whoever answered the email fastest. Then they discover, three months in, that 30 percent of delivered hours fail their own QA, the cameras were never time-synced to spec, and the contract says nothing about redelivery.

Our thesis, argued with numbers throughout this guide: buying robot data is buying a vendor’s pipeline, and because the defects that matter are latent, only three instruments filter them before signature: a written acceptance spec, a weighted scorecard, and a paid pilot. This guide gives you that procurement discipline, which the category is missing. You will get a 10-criterion weighted vendor scorecard, a 15-question RFP bank grouped by category, a red flags list built from real deals, and a 50-hour paid pilot protocol with pass/fail metrics. Everything here is designed to be copied into your next vendor evaluation.

We run teleoperation and egocentric capture pipelines at DexSet, which means we sit on the receiving end of RFPs every week. We have seen which questions separate serious buyers from tourists, and we have watched deals go sideways when those questions were never asked. This playbook is the document we wish every buyer sent us.

TL;DR: Key Takeaways

  • A robotics data buyer’s playbook is a structured process (RFP, weighted scorecard, red flags, paid pilot) for selecting and managing robot training data vendors.
  • Score vendors on 10 weighted criteria. Data quality SLAs, modality coverage, and calibration/sync spec carry the most weight (39 of 100 points combined).
  • Never sign an annual contract without a paid pilot. Our recommended protocol: 50 hours, measure usable-hour yield (target 85 percent or higher) and policy success delta on fixed eval tasks.
  • Pricing opacity is a signal, not an inconvenience. Vendors who publish per-hour ranges tend to survive QA audits; vendors who quote only after “discovery calls” often cannot.
  • Require delivery in a standard format (LeRobot, RLDS, or HDF5 with documented schema). A proprietary-format-only vendor is a lock-in risk.

What Is a Robotics Data Buyer’s Playbook?

A robotics data buyer’s playbook is a repeatable procurement process for evaluating, piloting, and contracting robot training data vendors, built around four artifacts: a request for proposal (RFP), a weighted vendor scorecard, a red flags checklist, and a pilot evaluation protocol. The playbook exists because robot demonstration data is a specification-heavy purchase, closer to contract manufacturing than to SaaS: what you receive is physical work (teleoperation, human video capture, annotation) frozen into files, and defects are expensive to detect after the fact.

The four artifacts map to four decisions:

  • RFP: which vendors are worth talking to at all.
  • Scorecard: how to compare the ones who respond, on the same axes, with weights that reflect your actual risk.
  • Red flags: which vendors to eliminate regardless of score.
  • Pilot protocol: whether the winning vendor’s production output matches their sample, before you commit annual budget.

Entity chain, stated plainly: your VLA model (OpenVLA, π0, GR00T-class) is trained by imitation learning on demonstration episodes; those episodes come from teleoperation rigs (ALOHA-style bimanual arms, VR-based systems) or egocentric human capture; the vendor’s calibration, sync, and QA pipeline determines whether those episodes are usable; your procurement process determines which vendor pipeline you inherit. Buying data is buying a pipeline.

Why Robot Training Data Is a Different Kind of Purchase

Robot training data procurement differs from every other data purchase because quality is defined by physics, not labels. In text or image annotation, a bad label is visible on screen. In robot data, the defects that kill model performance are invisible in a preview: a 40 ms sync offset between camera frames and joint states, extrinsics that drifted after someone bumped a wrist camera, gripper actions clipped at the edge of the calibration range. The episode looks fine. The policy trained on it does not.

Three properties follow from this:

Defects are latent. You often cannot see them until you train. This is why the pilot protocol below trains an actual policy instead of just eyeballing playback.

Specs are multidimensional. A single “hour of data” bundles camera count and placement (egocentric, exocentric, wrist), mono versus stereo, resolution and fps, depth, proprioception rate, action space, task diversity, and annotation depth. Two vendors quoting the same price are almost never quoting the same product.

Formats determine integration cost. The ecosystem has converged on a handful of formats: the LeRobot dataset format from Hugging Face (github.com/huggingface/lerobot), RLDS/TFDS episodes used by Open X-Embodiment (github.com/google-research/rlds; arxiv.org/abs/2310.08864), raw HDF5 in ALOHA conventions (arxiv.org/abs/2304.13705), and log formats like MCAP (github.com/foxglove/mcap) and rosbag2 (github.com/ros2/rosbag2) for ROS 2 stacks (docs.ros.org). A vendor who cannot deliver in at least one of these adds weeks of conversion work to every delivery.

Here is how the common delivery formats compare from a buyer’s perspective:

Format Origin / ecosystem Best for Random access Tooling maturity Buyer risk if this is the only option
LeRobot v2.x Hugging Face LeRobot Training loops, HF hub distribution Good (parquet + mp4) High, active community Low
RLDS / TFDS Google, Open X-Embodiment TF/JAX pipelines, OXE mixing Good High in TF world, thinner in PyTorch Low to medium
HDF5 (ALOHA-style) ACT / ALOHA papers Bimanual teleop episodes Good Medium, schema varies by lab Medium (demand a schema doc)
MCAP Foxglove Multimodal logging, ROS 2 Good Growing fast Medium
rosbag2 ROS 2 Robot-native logging Moderate High in ROS, low outside Medium
Proprietary format Single vendor Nothing you control Unknown None outside vendor High: lock-in, conversion cost, audit difficulty

The 10-Criterion Weighted Vendor Scorecard

The DexSet vendor scorecard evaluates a robotics data provider on ten criteria, each scored 1 to 5 and multiplied by a weight, for a maximum of 500 weighted points. The weights below reflect where we see deals actually fail; adjust them to your program, but keep quality, modality, and calibration at the top.

# Criterion Weight What a 5 looks like What a 1 looks like
1 Data quality SLAs 15 Contractual usable-hour yield (e.g., 90 percent pass your acceptance spec), free redelivery of failures “We do QA” with no numbers
2 Modality coverage 12 Ego + exo + wrist cams, stereo option, depth, synced proprioception, tactile roadmap Single exocentric mono camera
3 Calibration & sync spec 12 Published intrinsics/extrinsics per rig, time sync under 10 ms across streams, drift checks per shift “Cameras are calibrated” with no numbers or recheck cadence
4 Annotation accuracy 10 Documented rubric, dual-pass or audit sampling, 97 percent or higher on independent audit Unversioned labels, no audit trail
5 Throughput (hours/week) 10 Stated capacity per rig type, evidence of sustained delivery at your volume “We can scale to anything” without rig counts
6 Pricing transparency 10 Published per-hour ranges by rig and task complexity Pricing only after discovery calls
7 Privacy & consent chain 8 Signed operator/participant consent on file, face/bystander handling policy, GDPR posture No documented consent process
8 Format compatibility 8 Native LeRobot and RLDS export, HDF5/MCAP on request, schema docs Proprietary format only
9 Pilot terms 8 Offers paid pilot with buyer-defined acceptance criteria Refuses pilots or offers only cherry-picked samples
10 IP & licensing 7 You own delivered data and models trained on it; no reuse without consent Vendor retains broad reuse rights, ambiguous model ownership

How to use it: have two people score independently from the RFP responses and sample data, then reconcile. Anything below 350/500 exits the process. Anything scoring 1 or 2 on criteria 1, 3, or 10 exits regardless of total, because quality SLAs, calibration, and licensing are the three areas where a weak answer becomes an unrecoverable problem after signature.

The RFP Question Bank: 15 Questions That Do the Work

A robotics data RFP is a short, specific document (five pages beats fifty) whose questions force vendors to commit to numbers. Below are 15 questions we recommend, grouped by category. Vendors who answer all 15 with specifics belong on your shortlist. Vendors who answer with adjectives do not.

Quality and QA 1. What usable-hour yield do you commit to contractually against an acceptance spec we define, and what happens to hours that fail (redelivery, credit, or refund)? 2. Describe your per-episode QA pipeline: what is checked automatically, what is checked by humans, and what percentage of episodes get a second-pass review? 3. Provide annotation accuracy from your most recent independent audit, and the rubric it was measured against.

Hardware, calibration, and sync 4. For each rig type you would use on our program, list cameras (placement, mono/stereo, resolution, fps), depth sensors, and proprioception rates. 5. What is your maximum time-sync error across camera, joint-state, and action streams, and how is it measured and rechecked during production? 6. How often are intrinsics and extrinsics recalibrated, and do delivered episodes include per-rig calibration files?

Operations and throughput 7. How many rigs and operators would be dedicated to our program, and what sustained hours per week does that support at our spec? 8. What was your actual delivered volume for your largest program in the last six months (hours, not episodes)? 9. What is your process when task success rates drop or instructions change mid-program?

Commercial 10. Provide your per-hour pricing range by rig type and task complexity, and state every cost that is not included in it (setup, annotation, redelivery, format conversion). 11. What are your paid pilot terms: minimum hours, price, timeline, and whether buyer-defined acceptance criteria apply? 12. In which formats do you deliver natively (LeRobot, RLDS, HDF5, MCAP, rosbag2), and can you share a schema document today?

Legal and privacy 13. Who owns the delivered data and any models trained on it, and do you retain any right to reuse our episodes for other customers or your own models? 14. Describe your consent chain: what operators and any captured bystanders sign, and how consent records map to delivered episodes. 15. What happens to our task definitions, environment setups, and prompts after the engagement ends?

Red Flags: When to Walk Away

A red flag in robotics data procurement is any vendor behavior that predicts unrecoverable problems after contract signature. These are the ones we treat as disqualifying, based on deals we have watched from both sides:

  • Pricing available only after a discovery call. Opacity here usually means price is set by your budget, not their costs.
  • Refusal of a paid pilot with buyer-defined acceptance criteria. A vendor confident in their pipeline will take your money to prove it.
  • Sample data that is not from a production rig. Ask directly. Showcase rigs and production rigs can be different machines run by different people.
  • No stated time-sync tolerance. If they have never measured it, your training team will be the first to.
  • Proprietary delivery format with no export path. Lock-in plus audit difficulty in one package.
  • No consent documentation for operators or captured humans. This becomes your legal problem, not theirs.
  • Broad data reuse rights buried in the MSA. Your task distribution is competitive information.
  • “Unlimited scale” claims without rig counts. Throughput is rigs times shifts times yield. Anyone who will not show the multiplication is guessing.
  • No redelivery or credit policy for failed hours. QA without consequences is marketing.

Cost and Economics: What Robot Data Actually Costs

Robot training data pricing in 2026 clusters into ranges that depend on rig type, task complexity, and annotation depth, and any playbook needs those ranges to sanity-check quotes. The figures below are DexSet benchmarks and typical ranges we see across the market; treat them as calibration, not quotes.

Data type Typical market range (per hour) Main cost drivers Typical usable-hour yield we see
Bimanual teleoperation (ALOHA-class rig) $28–$60 Operator skill, task resets, rig count 80–90 percent
Humanoid on-robot teleoperation $150–$400 Robot cost and uptime, safety oversight, pilot skill 70–85 percent
Egocentric human video (mono to stereo + IMU) $15–$40 Participant recruiting, consent, headset hardware 80–90 percent
Exocentric multi-view capture $20–$50 Camera count, calibration, studio setup 80–90 percent
Dense annotation add-on +$8–$25 Label depth, audit sampling n/a
Independent QA add-on +$5–$12 Sampling rate, audit depth n/a

Two economics rules matter more than the sticker price:

Rule 1: price per usable hour, not per delivered hour. A $35/hr vendor at 70 percent yield costs you $50 per usable hour. A $44/hr vendor at 90 percent yield costs $48.89, arrives with less rework, and does not poison your training mix with borderline episodes.

Rule 2: model the pipeline, not the purchase. A 1,000-hour teleop program at $40/hr is $40,000 in data, but plan another 15 to 25 percent for integration, conversion, storage, and your own audit time if the vendor scores poorly on format compatibility and QA. Vendors who score 4 to 5 on criteria 1, 4, and 8 compress that overhead, which is usually worth more than a $5/hr discount.

The 50-Hour Paid Pilot Protocol

A pilot evaluation protocol is a fixed, paid, pre-contract engagement that measures whether a vendor’s production pipeline meets your acceptance spec, using metrics you define before the first hour is captured. We recommend 50 hours: large enough to expose process problems, small enough that a failed pilot costs weeks rather than quarters.

Protocol:

  • Fix the spec first. Write the acceptance spec (camera config, sync tolerance, task list, annotation rubric, delivery format) before contacting vendors. The pilot tests the vendor against the spec, not the spec against the vendor.
  • Pay for it. 50 hours at market rates is $1,500 to $4,500 for most modalities. Paying keeps the vendor’s incentives honest and gets you production treatment, not showcase treatment.
  • Demand production conditions. Same rigs, same operators, same QA pipeline that would run your annual contract. Put this in the pilot agreement.
  • Measure usable-hour yield. Run every delivered episode through your acceptance checks. Target: 85 percent or higher for teleop, 90 percent or higher for egocentric video.
  • Audit annotations on a sample. Independently re-label 5 percent of episodes. Target: 97 percent agreement or higher.
  • Train a policy and measure the delta. Fine-tune a fixed baseline (an ACT or diffusion policy head, or a small VLA) on your existing data alone, then on existing data plus pilot data. Evaluate both on the same fixed task set. The pilot passes only if success rate improves; flat or negative delta at 50 hours predicts flat or negative at 1,000.
  • Test the loader. Delivered data should load into your LeRobot or RLDS pipeline within one engineer-day. Log every schema surprise; each one recurs at scale.
  • Decide on numbers. Yield, audit accuracy, policy delta, loader time. Four numbers, agreed in advance, written into the pilot agreement.

Case Study: A Humanoid Foundation Model Team Rebuilds Its Buying Process

One humanoid foundation model team we work with came to us after a year of ad hoc purchasing: three vendors, three formats, no shared acceptance spec, and a training team spending roughly a day per delivery writing conversion scripts and triaging bad episodes. Their internal estimate was that a quarter of purchased hours never reached a training run.

They rebuilt procurement around the artifacts in this guide. The spec came first: stereo egocentric plus wrist cameras, sync under 10 ms, LeRobot delivery, a 60-task manipulation list. The RFP went to six vendors; four answered with numbers, two answered with adjectives and were dropped. The scorecard separated the four finalists cleanly, mostly on quality SLAs and licensing terms, where two vendors wanted broad reuse rights. Both finalists ran 50-hour paid pilots. One delivered 90 percent usable-hour yield and a measurable success-rate improvement on the fixed eval set; the other delivered 78 percent yield and failed the loader test.

The team signed with the first vendor. The numbers that mattered afterward: usable-hour yield across the first 1,000 contracted hours held within two points of the pilot, and data engineering time per delivery dropped from about a day to under an hour because the format and schema were locked before signature. No named clients, no invented revenue figures, just the pattern we see repeatedly: the pilot predicted production, because the pilot was run under production conditions.

Get the Template

The fastest way to apply this playbook is to not rebuild it. We packaged the full RFP question bank (the 15 questions above plus 20 more), the weighted scorecard as a spreadsheet with the math built in, the red flags checklist, and the pilot agreement language into a single download. Send the RFP as-is or strip it to the sections that match your program.

Related reading on this site: why procurement is the real bottleneck in physical AI, comparing data sourcing approaches and their costs, a humanoid team’s VLA data procurement case study, and five hidden challenges in robotics data RFPs.

Put the Playbook to Work

Ready to scope your data program? Talk to our team.

Frequently Asked Questions

What is a robotics data buyer’s playbook?

A robotics data buyer’s playbook is a structured procurement process for robot training data, built around four artifacts: an RFP with specification-forcing questions, a weighted vendor scorecard, a red flags checklist, and a paid pilot protocol with predefined pass/fail metrics.

Five to eight. Fewer gives you no basis for comparison; more than eight means your spec is probably too vague to have filtered anyone out, and evaluation time grows linearly with responses.

In our benchmarks, mature teleoperation pipelines deliver 80 to 90 percent of hours passing a well-defined acceptance spec. Below 80 percent, rework and training-mix contamination usually erase any price advantage.

Paid. A paid 50-hour pilot ($1,500 to $4,500 at typical market rates) obligates the vendor to run production processes and lets you enforce buyer-defined acceptance criteria. Free samples are curated by definition.

Require native delivery in at least one of LeRobot, RLDS/TFDS, or documented HDF5, with MCAP or rosbag2 as options for ROS 2 stacks. Reject proprietary-only formats; they create lock-in and make independent audits harder.

Start from the DexSet weights (quality SLAs 15, modality coverage 12, calibration/sync 12) and shift weight toward whichever criterion caused your last data problem. Keep quality, calibration, and licensing as automatic-disqualification criteria regardless of weights.

Download the RFP Template + Vendor Scorecard (XLSX) and send it to your shortlist this week, or book a demo and we will walk you through how DexSet answers all 15 questions, numbers included.

5 Hidden Challenges in Robotics Data RFPs (and How to Solve Them)

The contract was one signature away when the buyer’s counsel asked who owned models trained on the delivered episodes, and the call went quiet for a long moment. Nobody had put the question in the RFP. The clause that eventually surfaced in the vendor’s MSA granted reuse rights over “aggregated and derived data”, which is to say, over the buyer’s task distribution. That deal survived. The assumption that the important questions were already in the document did not.

The obvious RFP challenges are the ones buyers plan for: price, volume, timeline. Those rarely sink a data program. What sinks programs are the questions that never make it into the document, because the buyer did not know the failure mode existed until it arrived inside a delivery.

I run teleoperation rigs for a living, which means I also read the RFPs that reach us. The pattern is consistent: the sections buyers write carefully (pricing tables, volume schedules) cover the risks they have already survived, and the sections they skip (sync tolerances, consent chains, schema versioning) cover the risks they have not met yet. An RFP is a map of its author’s scar tissue.

This post covers the five challenges we most often see missing, why each one stays hidden until it costs money, and the specific question or clause that fixes it. All five fixes are built into the templates in the Robotics Data Buyer’s Playbook.

Key Takeaways

  • The five hidden RFP challenges: undefined sync and calibration specs, licensing and reuse traps, throughput fiction, missing consent chains, and format drift between deliveries.
  • Each has a one-question or one-clause fix that costs nothing to include and is expensive to omit.
  • The common thread: RFPs fail on what they do not ask. Vendors answer the questions in the document, not the ones you meant.

Challenge 1: The Sync Spec That Nobody Defines

A sync spec is the maximum allowed time offset between data streams (cameras, joint states, actions) in a delivered episode, and most RFPs never state one. The omission stays hidden because unsynced data plays back fine to a human eye. It surfaces months later as an imitation learning policy that plateaus below expectations, because the state-action pairs it learned from were misaligned by tens of milliseconds.

The fix: state a tolerance and demand measurement. Our recommended RFP language: “Maximum time-sync error across all streams shall be 10 ms; vendor shall describe how sync is measured, how often it is rechecked during production, and how violations are flagged in delivered metadata.” A vendor who has never measured their sync error will reveal it in one sentence, which is exactly what you paid the question to find out. The same applies to calibration: require per-rig intrinsics and extrinsics files with every delivery and a stated recalibration cadence.

Challenge 2: The Licensing Trap in the MSA

The licensing trap is contract language granting the vendor rights to reuse your delivered episodes, your task definitions, or data derived from them for other customers or their own models. It hides in the master service agreement rather than the RFP response, usually as one clause about “aggregated or derived data.” Your task list and environment setups are competitive information; a reuse clause quietly donates them to whoever buys from the same vendor next.

The fix: ask it in the RFP, before the MSA exists: “Who owns delivered data and models trained on it? Do you retain any right to reuse our episodes, task definitions, or derivatives for any other purpose?” Then score the answer on your vendor scorecard and treat broad reuse rights as disqualifying. Negotiating this after signature means bargaining from a position you already gave away.

Challenge 3: Throughput Fiction

Throughput fiction is a capacity claim (“we can scale to any volume”) unsupported by the arithmetic that produces real hours: rig count times shifts times operators times usable-hour yield. It stays hidden because it fails late, four or six weeks into a ramp, when a vendor sized for 40 hours a week is contractually committed to 150 and quality starts absorbing the difference.

The fix: make the multiplication mandatory. RFP question: “For our program spec, state the number of rigs and operators you would dedicate, the sustained hours per week that supports, and your actual delivered volume for your largest program in the last six months.” Any credible operation knows these numbers instantly. On teleoperation specifically, sustained real-world throughput per bimanual rig runs far below the naive shift math once you account for resets, calibration checks, operator breaks, and QA rejections; in our own operations, a single ALOHA-class rig sustains roughly 20 to 30 usable hours per week, not the 60 to 80 the shift calendar implies.

Challenge 4: The Missing Consent Chain

A consent chain is the documented link between every delivered episode and the signed consent of every human in it: teleoperators, egocentric camera wearers, and bystanders captured in frame. RFPs skip it because it feels like legal boilerplate. It becomes real the day your model ships in a product and your counsel asks whether the training data included identifiable people who never agreed to it, a question with GDPR and biometric-law consequences you cannot retroactively fix.

The fix: two RFP questions. “Describe what operators and captured participants sign, and how consent records map to delivered episodes.” And: “What is your process for faces and bystanders in egocentric capture: avoidance, blurring, or documented consent?” Public egocentric corpora set the reference point here; Ego4D and EgoExo4D shipped with documented consent and privacy processes (arxiv.org/abs/2110.07058; arxiv.org/abs/2311.18259), and a commercial vendor should clear the bar academic datasets already cleared.

Challenge 5: Format Drift Between Deliveries

Format drift is schema change between deliveries from the same vendor: a renamed key in an HDF5 file, a reordered camera list, a new compression setting, an action space silently rescaled. Single-delivery evaluation cannot catch it by definition, which is why it hides through every pilot and surfaces as a broken training pipeline at 2 a.m. before a deadline. The ecosystem’s convergence on versioned formats (LeRobot’s dataset versions, RLDS episode specs, MCAP channel schemas; github.com/huggingface/lerobot, github.com/google-research/rlds, github.com/foxglove/mcap) exists precisely because ad hoc schemas drift.

The fix: contract the schema, not just the format. Require a written schema document as an RFP deliverable, version it as a contract appendix, and add the clause: “Schema changes require written notice one delivery cycle in advance; unannounced schema changes constitute delivery failure.” Then validate mechanically: a loader script in CI that runs on every delivery costs a day to write and catches drift while it is still the vendor’s problem.

Why These Five Stay Hidden

The common mechanism behind all five challenges is delayed feedback: each failure surfaces weeks or months after the decision that caused it, in a different team’s backlog. Sync defects appear as ML debugging tickets. Licensing traps appear in legal review of a partnership, a year later. Throughput fiction appears as a slipped training milestone that gets blamed on the schedule, not the contract. Because the pain lands far from the RFP, the RFP never learns. The fix is not vigilance, which does not scale; it is putting the five questions into a template so they get asked by default, on every deal, including the ones that feel too small or too friendly to need them. Friendly deals with partner labs are where we see the format-drift and consent gaps most often, precisely because nobody wanted to send paperwork to a friend.

The Five Challenges at a Glance

Hidden challenge Why it stays hidden Cost when it surfaces The one-line fix
Undefined sync/calibration spec Bad sync looks fine in playback Weeks of training debugging Require ≤10 ms tolerance, measured and rechecked
Licensing reuse trap Lives in the MSA, not the RFP Task distribution leaks to competitors Ask ownership and reuse questions in the RFP; disqualify broad reuse
Throughput fiction Fails weeks into ramp, not at signing Missed training milestones Demand rigs × shifts × yield arithmetic plus 6-month delivery history
Missing consent chain Feels like boilerplate until launch Legal exposure you cannot backfill Require consent-to-episode mapping and a bystander policy
Format drift Invisible in any single delivery Broken pipelines, silent data corruption Version the schema in the contract; validate every delivery in CI

Close the Gaps Before You Send

Ready to scope your data program? Talk to our team.

Frequently Asked Questions

What is the most expensive hidden challenge in robotics data RFPs?

Format drift and undefined sync specs compete for the title. Sync problems corrupt what the model learns; format drift breaks the pipeline that feeds it. Both are cheap to prevent with one RFP requirement and expensive to diagnose after delivery.

Ask for the arithmetic: dedicated rigs, operators, shifts, and usable-hour yield, plus actual delivered volume on their largest recent program. As a sanity check, a bimanual teleop rig sustains roughly 20 to 30 usable hours per week in real operations.

Any retained right to reuse your episodes, task definitions, or derivatives for other customers or their own models. Ownership of delivered data and models trained on it should sit with the buyer, stated in the RFP response before the MSA stage.

Yes. Teleoperators are identifiable humans generating biometric-adjacent data, and egocentric capture routinely includes bystanders. Require consent records that map to delivered episodes, matching the standard public datasets like Ego4D already meet.