Skip to main content

Dexset

The Robotics Data Buyer’s Playbook: RFPs, Scorecards, and Pilots That Actually Work (2026 Guide)

QA_FAIL episode_0847.hdf5: sync_offset_max 41ms (cam_high vs qpos). That log line, or one very like it, is how many robot data programs discover they have a procurement problem: the flag fires weeks after signature, the contract never defined a sync tolerance, and the vendor technically delivered exactly what was ordered. The team can tell you its GPU budget to the dollar. It cannot tell you which of three vendors quoting $30, $45, and $80 per teleoperation hour would have caught that flag before shipping.

That gap exists for a structural reason. Robot training data has no established procurement discipline. Software buyers inherited decades of RFP practice from enterprise IT. Robotics data buyers inherited nothing, because the category barely existed before large-scale imitation learning and VLA models made demonstration data a line item worth millions. So teams improvise: they buy on price, on a demo video, or on whoever answered the email fastest. Then they discover, three months in, that 30 percent of delivered hours fail their own QA, the cameras were never time-synced to spec, and the contract says nothing about redelivery.

Our thesis, argued with numbers throughout this guide: buying robot data is buying a vendor’s pipeline, and because the defects that matter are latent, only three instruments filter them before signature: a written acceptance spec, a weighted scorecard, and a paid pilot. This guide gives you that procurement discipline, which the category is missing. You will get a 10-criterion weighted vendor scorecard, a 15-question RFP bank grouped by category, a red flags list built from real deals, and a 50-hour paid pilot protocol with pass/fail metrics. Everything here is designed to be copied into your next vendor evaluation.

We run teleoperation and egocentric capture pipelines at DexSet, which means we sit on the receiving end of RFPs every week. We have seen which questions separate serious buyers from tourists, and we have watched deals go sideways when those questions were never asked. This playbook is the document we wish every buyer sent us.

TL;DR: Key Takeaways

  • A robotics data buyer’s playbook is a structured process (RFP, weighted scorecard, red flags, paid pilot) for selecting and managing robot training data vendors.
  • Score vendors on 10 weighted criteria. Data quality SLAs, modality coverage, and calibration/sync spec carry the most weight (39 of 100 points combined).
  • Never sign an annual contract without a paid pilot. Our recommended protocol: 50 hours, measure usable-hour yield (target 85 percent or higher) and policy success delta on fixed eval tasks.
  • Pricing opacity is a signal, not an inconvenience. Vendors who publish per-hour ranges tend to survive QA audits; vendors who quote only after “discovery calls” often cannot.
  • Require delivery in a standard format (LeRobot, RLDS, or HDF5 with documented schema). A proprietary-format-only vendor is a lock-in risk.

What Is a Robotics Data Buyer’s Playbook?

A robotics data buyer’s playbook is a repeatable procurement process for evaluating, piloting, and contracting robot training data vendors, built around four artifacts: a request for proposal (RFP), a weighted vendor scorecard, a red flags checklist, and a pilot evaluation protocol. The playbook exists because robot demonstration data is a specification-heavy purchase, closer to contract manufacturing than to SaaS: what you receive is physical work (teleoperation, human video capture, annotation) frozen into files, and defects are expensive to detect after the fact.

The four artifacts map to four decisions:

  • RFP: which vendors are worth talking to at all.
  • Scorecard: how to compare the ones who respond, on the same axes, with weights that reflect your actual risk.
  • Red flags: which vendors to eliminate regardless of score.
  • Pilot protocol: whether the winning vendor’s production output matches their sample, before you commit annual budget.

Entity chain, stated plainly: your VLA model (OpenVLA, π0, GR00T-class) is trained by imitation learning on demonstration episodes; those episodes come from teleoperation rigs (ALOHA-style bimanual arms, VR-based systems) or egocentric human capture; the vendor’s calibration, sync, and QA pipeline determines whether those episodes are usable; your procurement process determines which vendor pipeline you inherit. Buying data is buying a pipeline.

Why Robot Training Data Is a Different Kind of Purchase

Robot training data procurement differs from every other data purchase because quality is defined by physics, not labels. In text or image annotation, a bad label is visible on screen. In robot data, the defects that kill model performance are invisible in a preview: a 40 ms sync offset between camera frames and joint states, extrinsics that drifted after someone bumped a wrist camera, gripper actions clipped at the edge of the calibration range. The episode looks fine. The policy trained on it does not.

Three properties follow from this:

Defects are latent. You often cannot see them until you train. This is why the pilot protocol below trains an actual policy instead of just eyeballing playback.

Specs are multidimensional. A single “hour of data” bundles camera count and placement (egocentric, exocentric, wrist), mono versus stereo, resolution and fps, depth, proprioception rate, action space, task diversity, and annotation depth. Two vendors quoting the same price are almost never quoting the same product.

Formats determine integration cost. The ecosystem has converged on a handful of formats: the LeRobot dataset format from Hugging Face (github.com/huggingface/lerobot), RLDS/TFDS episodes used by Open X-Embodiment (github.com/google-research/rlds; arxiv.org/abs/2310.08864), raw HDF5 in ALOHA conventions (arxiv.org/abs/2304.13705), and log formats like MCAP (github.com/foxglove/mcap) and rosbag2 (github.com/ros2/rosbag2) for ROS 2 stacks (docs.ros.org). A vendor who cannot deliver in at least one of these adds weeks of conversion work to every delivery.

Here is how the common delivery formats compare from a buyer’s perspective:

Format Origin / ecosystem Best for Random access Tooling maturity Buyer risk if this is the only option
LeRobot v2.x Hugging Face LeRobot Training loops, HF hub distribution Good (parquet + mp4) High, active community Low
RLDS / TFDS Google, Open X-Embodiment TF/JAX pipelines, OXE mixing Good High in TF world, thinner in PyTorch Low to medium
HDF5 (ALOHA-style) ACT / ALOHA papers Bimanual teleop episodes Good Medium, schema varies by lab Medium (demand a schema doc)
MCAP Foxglove Multimodal logging, ROS 2 Good Growing fast Medium
rosbag2 ROS 2 Robot-native logging Moderate High in ROS, low outside Medium
Proprietary format Single vendor Nothing you control Unknown None outside vendor High: lock-in, conversion cost, audit difficulty

The 10-Criterion Weighted Vendor Scorecard

The DexSet vendor scorecard evaluates a robotics data provider on ten criteria, each scored 1 to 5 and multiplied by a weight, for a maximum of 500 weighted points. The weights below reflect where we see deals actually fail; adjust them to your program, but keep quality, modality, and calibration at the top.

# Criterion Weight What a 5 looks like What a 1 looks like
1 Data quality SLAs 15 Contractual usable-hour yield (e.g., 90 percent pass your acceptance spec), free redelivery of failures “We do QA” with no numbers
2 Modality coverage 12 Ego + exo + wrist cams, stereo option, depth, synced proprioception, tactile roadmap Single exocentric mono camera
3 Calibration & sync spec 12 Published intrinsics/extrinsics per rig, time sync under 10 ms across streams, drift checks per shift “Cameras are calibrated” with no numbers or recheck cadence
4 Annotation accuracy 10 Documented rubric, dual-pass or audit sampling, 97 percent or higher on independent audit Unversioned labels, no audit trail
5 Throughput (hours/week) 10 Stated capacity per rig type, evidence of sustained delivery at your volume “We can scale to anything” without rig counts
6 Pricing transparency 10 Published per-hour ranges by rig and task complexity Pricing only after discovery calls
7 Privacy & consent chain 8 Signed operator/participant consent on file, face/bystander handling policy, GDPR posture No documented consent process
8 Format compatibility 8 Native LeRobot and RLDS export, HDF5/MCAP on request, schema docs Proprietary format only
9 Pilot terms 8 Offers paid pilot with buyer-defined acceptance criteria Refuses pilots or offers only cherry-picked samples
10 IP & licensing 7 You own delivered data and models trained on it; no reuse without consent Vendor retains broad reuse rights, ambiguous model ownership

How to use it: have two people score independently from the RFP responses and sample data, then reconcile. Anything below 350/500 exits the process. Anything scoring 1 or 2 on criteria 1, 3, or 10 exits regardless of total, because quality SLAs, calibration, and licensing are the three areas where a weak answer becomes an unrecoverable problem after signature.

The RFP Question Bank: 15 Questions That Do the Work

A robotics data RFP is a short, specific document (five pages beats fifty) whose questions force vendors to commit to numbers. Below are 15 questions we recommend, grouped by category. Vendors who answer all 15 with specifics belong on your shortlist. Vendors who answer with adjectives do not.

Quality and QA 1. What usable-hour yield do you commit to contractually against an acceptance spec we define, and what happens to hours that fail (redelivery, credit, or refund)? 2. Describe your per-episode QA pipeline: what is checked automatically, what is checked by humans, and what percentage of episodes get a second-pass review? 3. Provide annotation accuracy from your most recent independent audit, and the rubric it was measured against.

Hardware, calibration, and sync 4. For each rig type you would use on our program, list cameras (placement, mono/stereo, resolution, fps), depth sensors, and proprioception rates. 5. What is your maximum time-sync error across camera, joint-state, and action streams, and how is it measured and rechecked during production? 6. How often are intrinsics and extrinsics recalibrated, and do delivered episodes include per-rig calibration files?

Operations and throughput 7. How many rigs and operators would be dedicated to our program, and what sustained hours per week does that support at our spec? 8. What was your actual delivered volume for your largest program in the last six months (hours, not episodes)? 9. What is your process when task success rates drop or instructions change mid-program?

Commercial 10. Provide your per-hour pricing range by rig type and task complexity, and state every cost that is not included in it (setup, annotation, redelivery, format conversion). 11. What are your paid pilot terms: minimum hours, price, timeline, and whether buyer-defined acceptance criteria apply? 12. In which formats do you deliver natively (LeRobot, RLDS, HDF5, MCAP, rosbag2), and can you share a schema document today?

Legal and privacy 13. Who owns the delivered data and any models trained on it, and do you retain any right to reuse our episodes for other customers or your own models? 14. Describe your consent chain: what operators and any captured bystanders sign, and how consent records map to delivered episodes. 15. What happens to our task definitions, environment setups, and prompts after the engagement ends?

Red Flags: When to Walk Away

A red flag in robotics data procurement is any vendor behavior that predicts unrecoverable problems after contract signature. These are the ones we treat as disqualifying, based on deals we have watched from both sides:

  • Pricing available only after a discovery call. Opacity here usually means price is set by your budget, not their costs.
  • Refusal of a paid pilot with buyer-defined acceptance criteria. A vendor confident in their pipeline will take your money to prove it.
  • Sample data that is not from a production rig. Ask directly. Showcase rigs and production rigs can be different machines run by different people.
  • No stated time-sync tolerance. If they have never measured it, your training team will be the first to.
  • Proprietary delivery format with no export path. Lock-in plus audit difficulty in one package.
  • No consent documentation for operators or captured humans. This becomes your legal problem, not theirs.
  • Broad data reuse rights buried in the MSA. Your task distribution is competitive information.
  • “Unlimited scale” claims without rig counts. Throughput is rigs times shifts times yield. Anyone who will not show the multiplication is guessing.
  • No redelivery or credit policy for failed hours. QA without consequences is marketing.

Cost and Economics: What Robot Data Actually Costs

Robot training data pricing in 2026 clusters into ranges that depend on rig type, task complexity, and annotation depth, and any playbook needs those ranges to sanity-check quotes. The figures below are DexSet benchmarks and typical ranges we see across the market; treat them as calibration, not quotes.

Data type Typical market range (per hour) Main cost drivers Typical usable-hour yield we see
Bimanual teleoperation (ALOHA-class rig) $28–$60 Operator skill, task resets, rig count 80–90 percent
Humanoid on-robot teleoperation $150–$400 Robot cost and uptime, safety oversight, pilot skill 70–85 percent
Egocentric human video (mono to stereo + IMU) $15–$40 Participant recruiting, consent, headset hardware 80–90 percent
Exocentric multi-view capture $20–$50 Camera count, calibration, studio setup 80–90 percent
Dense annotation add-on +$8–$25 Label depth, audit sampling n/a
Independent QA add-on +$5–$12 Sampling rate, audit depth n/a

Two economics rules matter more than the sticker price:

Rule 1: price per usable hour, not per delivered hour. A $35/hr vendor at 70 percent yield costs you $50 per usable hour. A $44/hr vendor at 90 percent yield costs $48.89, arrives with less rework, and does not poison your training mix with borderline episodes.

Rule 2: model the pipeline, not the purchase. A 1,000-hour teleop program at $40/hr is $40,000 in data, but plan another 15 to 25 percent for integration, conversion, storage, and your own audit time if the vendor scores poorly on format compatibility and QA. Vendors who score 4 to 5 on criteria 1, 4, and 8 compress that overhead, which is usually worth more than a $5/hr discount.

The 50-Hour Paid Pilot Protocol

A pilot evaluation protocol is a fixed, paid, pre-contract engagement that measures whether a vendor’s production pipeline meets your acceptance spec, using metrics you define before the first hour is captured. We recommend 50 hours: large enough to expose process problems, small enough that a failed pilot costs weeks rather than quarters.

Protocol:

  • Fix the spec first. Write the acceptance spec (camera config, sync tolerance, task list, annotation rubric, delivery format) before contacting vendors. The pilot tests the vendor against the spec, not the spec against the vendor.
  • Pay for it. 50 hours at market rates is $1,500 to $4,500 for most modalities. Paying keeps the vendor’s incentives honest and gets you production treatment, not showcase treatment.
  • Demand production conditions. Same rigs, same operators, same QA pipeline that would run your annual contract. Put this in the pilot agreement.
  • Measure usable-hour yield. Run every delivered episode through your acceptance checks. Target: 85 percent or higher for teleop, 90 percent or higher for egocentric video.
  • Audit annotations on a sample. Independently re-label 5 percent of episodes. Target: 97 percent agreement or higher.
  • Train a policy and measure the delta. Fine-tune a fixed baseline (an ACT or diffusion policy head, or a small VLA) on your existing data alone, then on existing data plus pilot data. Evaluate both on the same fixed task set. The pilot passes only if success rate improves; flat or negative delta at 50 hours predicts flat or negative at 1,000.
  • Test the loader. Delivered data should load into your LeRobot or RLDS pipeline within one engineer-day. Log every schema surprise; each one recurs at scale.
  • Decide on numbers. Yield, audit accuracy, policy delta, loader time. Four numbers, agreed in advance, written into the pilot agreement.

Case Study: A Humanoid Foundation Model Team Rebuilds Its Buying Process

One humanoid foundation model team we work with came to us after a year of ad hoc purchasing: three vendors, three formats, no shared acceptance spec, and a training team spending roughly a day per delivery writing conversion scripts and triaging bad episodes. Their internal estimate was that a quarter of purchased hours never reached a training run.

They rebuilt procurement around the artifacts in this guide. The spec came first: stereo egocentric plus wrist cameras, sync under 10 ms, LeRobot delivery, a 60-task manipulation list. The RFP went to six vendors; four answered with numbers, two answered with adjectives and were dropped. The scorecard separated the four finalists cleanly, mostly on quality SLAs and licensing terms, where two vendors wanted broad reuse rights. Both finalists ran 50-hour paid pilots. One delivered 90 percent usable-hour yield and a measurable success-rate improvement on the fixed eval set; the other delivered 78 percent yield and failed the loader test.

The team signed with the first vendor. The numbers that mattered afterward: usable-hour yield across the first 1,000 contracted hours held within two points of the pilot, and data engineering time per delivery dropped from about a day to under an hour because the format and schema were locked before signature. No named clients, no invented revenue figures, just the pattern we see repeatedly: the pilot predicted production, because the pilot was run under production conditions.

Get the Template

The fastest way to apply this playbook is to not rebuild it. We packaged the full RFP question bank (the 15 questions above plus 20 more), the weighted scorecard as a spreadsheet with the math built in, the red flags checklist, and the pilot agreement language into a single download. Send the RFP as-is or strip it to the sections that match your program.

Related reading on this site: why procurement is the real bottleneck in physical AI, comparing data sourcing approaches and their costs, a humanoid team’s VLA data procurement case study, and five hidden challenges in robotics data RFPs.

Put the Playbook to Work

Ready to scope your data program? Talk to our team.

Frequently Asked Questions

What is a robotics data buyer’s playbook?

A robotics data buyer’s playbook is a structured procurement process for robot training data, built around four artifacts: an RFP with specification-forcing questions, a weighted vendor scorecard, a red flags checklist, and a paid pilot protocol with predefined pass/fail metrics.

Five to eight. Fewer gives you no basis for comparison; more than eight means your spec is probably too vague to have filtered anyone out, and evaluation time grows linearly with responses.

In our benchmarks, mature teleoperation pipelines deliver 80 to 90 percent of hours passing a well-defined acceptance spec. Below 80 percent, rework and training-mix contamination usually erase any price advantage.

Paid. A paid 50-hour pilot ($1,500 to $4,500 at typical market rates) obligates the vendor to run production processes and lets you enforce buyer-defined acceptance criteria. Free samples are curated by definition.

Require native delivery in at least one of LeRobot, RLDS/TFDS, or documented HDF5, with MCAP or rosbag2 as options for ROS 2 stacks. Reject proprietary-only formats; they create lock-in and make independent audits harder.

Start from the DexSet weights (quality SLAs 15, modality coverage 12, calibration/sync 12) and shift weight toward whichever criterion caused your last data problem. Keep quality, calibration, and licensing as automatic-disqualification criteria regardless of weights.

Download the RFP Template + Vendor Scorecard (XLSX) and send it to your shortlist this week, or book a demo and we will walk you through how DexSet answers all 15 questions, numbers included.