Skip to main content

Dexset

Case Study: How We Scaled Privacy, Consent & Compliance in Human Data Collection for a VLA Model

Enoch Pakanati

This article is informational, not legal advice. The client is anonymized; no figures identify them.

Two teams went shopping for human demonstration data in the same period last year, and their paths split at the paperwork. One bought the cheapest hours available and deferred provenance to later; when counsel finally reviewed that corpus, the quote came back as months of retrofit review with an unknown write-off rate. The other, the humanoid foundation model team in this case study, made consent infrastructure a precondition of the program and got counsel approval of the corpus on first review. Same appetite for data, opposite outcomes.

The second team had already lived the first team’s story. Their previous vendor delivered hours cheaply but could not answer basic provenance questions: which consent version covered which episodes, whether Illinois participants had signed BIPA-standard written releases, what happened if someone withdrew. Their counsel froze procurement until those questions had answers. Data velocity had collided with evidence velocity. The new ask, heading into a funding-critical VLA training run: several thousand hours of egocentric and teleoperation data across household environments in three jurisdictions (US, EU, and one APAC market), delivered rolling, with documentation their lawyers and their own enterprise customers could audit.

This post argues one thesis with the numbers to back it: consent run as an engineering discipline is faster and cheaper than consent run as a cleanup project. We walk through how we structured the program, what it cost per stage, and where it nearly went wrong. If you are scoping a similar program, the point is not that you need us; it is that the mechanics below are reproducible by anyone willing to treat consent as infrastructure.

Key Takeaways – The program delivered several thousand hours of consented egocentric and teleop data across three jurisdictions with an episode-level audit trail from day one. – Consent infrastructure was built before capture: jurisdiction-specific templates, BIPA-standard written releases for US collection, digital signing, and a versioned consent registry. – PII redaction averaged roughly $4 per collected hour, at the low end of our $3 to $10 range, because scene sweeps kept bystander density down. – One participant withdrawal mid-program was resolved in days, not weeks, because every episode carried a consent ID. – Buyer counsel approved the corpus on first review, and provenance records plugged directly into the client’s EU AI Act documentation.

What the Program Had to Deliver

The requirement was a corpus a VLA training team could use and a legal team could defend, which in practice meant four deliverables per batch: episodes (egocentric video, stereo where specified, teleop trajectories with proprioception), redaction QA reports, dataset cards stating license and consent scope, and the audit trail linking each episode ID to a consent record. The client’s model stack sat in the standard lineage: teleoperation and egocentric demonstration data feeding imitation learning, training a VLA policy in the OpenVLA and pi0 tradition. Nothing exotic. The exotic part was the paperwork moving at the speed of the rigs.

Phase 1: Consent Infrastructure Before Cameras

Consent infrastructure means everything that must exist before the first participant signs: templates, signing flow, registry, and withdrawal mechanics. We spent the first three weeks here, capturing nothing.

Three jurisdictions meant three consent tracks. The EU track was drafted to the explicit consent standard, treating the data as potentially within GDPR Article 9 scope because downstream face embedding by the client could not be ruled out. The US track used written releases meeting the Illinois BIPA standard everywhere, not just for Illinois residents; standardizing on the strictest state was cheaper than gating participants by state of residence. The APAC track followed local counsel’s requirements. Every document was plain-language, named the modalities (video, audio, hand pose, gait, depth), named commercial model training and dataset licensing as uses, and specified the withdrawal channel. Each signed consent got an ID in a versioned registry. Total setup cost was a fixed five-figure line item; amortized over the program it added under $1.50 per hour.

Phase 2: Capture With Scene Sweeps Inline

A scene sweep is a pre-capture walkthrough that removes bystander and PII problems before the camera sees them, and it was the single highest-return control in the program. Household capture is where compliance usually leaks: family members wandering into frame, mail on counters, kids’ photos on refrigerators, phone screens face-up.

Every session began with a fifteen-minute sweep against a checklist: household members briefed and either consented or scheduled out, documents and mail cleared, screens down or masked, family photos turned. Sweeps were logged per session with the operator’s sign-off. The effect showed up directly in redaction economics; swept scenes came out of automated detection with a fraction of the flags of the unswept pilot sessions we ran for comparison. Capture itself ran on our standard rigs at our typical $28 to $60 per teleop hour depending on task complexity, with egocentric sessions in the same band.

Phase 3: Redaction as a Pipeline Stage, Not a Cleanup

Inline redaction means every captured session enters automated PII detection within a day and clears human QA within a week, while the program is still running and correctable. Faces and license plates went through automated detection and blurring (the open EgoBlur line of models is a reasonable public reference point for what this looks like), screens through a separate detector, and audio through a name-and-address transcript scrub.

We human-QA’d 15% of episodes, selected with a bias toward sessions flagged as higher risk by the sweep logs, and 100% of anything semi-public. Misses found in QA fed back into the sweep checklist within the same week; the recurring offenders were reflections (oven doors, microwaves) and phones picked up mid-task. Redaction averaged roughly $4 per collected hour across the program, at the low end of our published $3 to $10 range, and the sweeps are the reason.

The Withdrawal That Proved the System

A withdrawal request is the live-fire test of a consent chain, and we got one mid-program: a participant invoked their withdrawal right after several sessions had already shipped to the client. Under the program’s terms this meant halting further processing and deleting their episodes on both sides.

The registry resolved their consent ID to every affected episode in minutes. Deletion on our side, a certified deletion request to the client under the contract’s pass-through clause, and written confirmation closed out in days. The client’s Head of Data later told us this event, more than any deliverable, convinced their counsel the consent chain was real. A system that has never processed a withdrawal is a system you should assume does not work.

What the Program Cost, Stage by Stage

Program stageCost per collected hour (approx.)Share of total
Recruiting + consent workflow$1.50 to $2.504 to 6%
Capture (rig, operator, participant comp)$28 to $6075 to 85%
Scene sweepsunder $11 to 2%
PII redaction + QA~$4 average8 to 12%
Audit trail, dataset cards, deliveryunder $11 to 2%

Compliance, fully loaded, added roughly 15 to 20% to raw capture cost. The client’s alternative, a retrofit review of their previous vendor’s corpus, was quoted by their counsel as months of review with an unknown write-off rate. Against that baseline, 15 to 20% was not overhead; it was the cheapest item on the invoice.

What We Would Tell a Team Scoping This Today

Three lessons transfer to any program. First, build consent infrastructure before cameras roll; the three weeks we spent up front were repaid several times over in redaction cost and legal review speed. Second, standardize on the strictest applicable standard (explicit consent, BIPA-grade releases) instead of managing per-jurisdiction exceptions; uniformity is cheaper than cleverness. Third, treat the audit trail as a deliverable your buyer pays for, because that is exactly what it is at due-diligence time.

The full framework behind this program, including the six-link consent chain checklist, the de-identification cost table, and the buyer due-diligence question set, is in our pillar: The Complete Guide to Privacy, Consent & Compliance in Human Data Collection.

Next Step

Planning a VLA data program? Book a demo to see the audit trail from this case live, or start with the complete privacy, consent and compliance guide.

Frequently Asked Questions

How long does it take to set up compliant human data collection?

In this program, three weeks of consent infrastructure work (templates, signing flow, registry) preceded capture. With standing infrastructure already in place, comparable programs start capture within days.

Roughly 15 to 20% on top of raw capture: $1.50 to $2.50 per hour for consent workflow, about $4 per hour for redaction and QA, and small fixed costs for sweeps and audit trail tooling.

The consent registry resolves the participant’s consent ID to every affected episode, deletion runs on the vendor side, and a contractual pass-through clause obligates the buyer to delete their copies, with written confirmation closing the loop.

Standardizing on the strictest applicable US standard removes the need to gate participants by state and eliminates the risk of discovering an Illinois nexus after collection, at negligible extra cost per participant.

Per-batch redaction QA reports, dataset cards stating license terms and consent scope, and an audit trail linking every episode to a signed consent record. Those three artifacts are what buyer counsel and EU AI Act documentation both require.

Enoch Pakanati
Written by

Enoch Pakanati

Enoch Pakanati is the strategic architect behind DexSet’s mission to become the undisputed market leader in robotics training data. He oversees the company’s growth strategy, focusing on capturing dominant market share across all data modalities required for modern robotics, including egocentric capture, teleoperation, and simulation-to-real data pipelines.

At DexSet, Enoch is responsible for transforming the company’s deep technical capabilities into a market-leading brand that foundation model labs and robotics OEMs trust implicitly. He focuses on scaling DexSet’s global footprint and ensuring the company stays ahead of the industry’s rapidly evolving data needs. His leadership is centered on one objective: making DexSet the singular, global standard for the data that powers the robotics revolution.