Skip to main content

Dexset

Why Privacy, Consent & Compliance in Human Data Collection Is the Biggest Bottleneck in Physical AI

Enoch Pakanati

In January, we started logging idle time on our capture rigs across two collection programs, and the result was uncomfortable: the cameras spent more hours waiting than recording. Nothing was wrong with the hardware. GPUs waited too. The queue had formed in front of consent paperwork, redaction QA, and legal review, the stages nobody had instrumented because nobody thought of them as the pipeline.

That is this post’s thesis: the throughput of a physical AI data program is set by its compliance stages, not its capture capacity. The reason is structural. Physical AI cannot repeat the LLM playbook. Text models trained on data that already existed in public, while robot models need humans performing tasks on camera, mostly in private spaces, mostly generating biometric-adjacent signals like faces, voices, and body movement. Each hour has to be manufactured, and each manufactured hour carries obligations under GDPR, CCPA/CPRA, Illinois BIPA, and now the EU AI Act’s documentation duties. You can parallelize rigs. You cannot skip consent.

This post maps where the throughput actually dies: which stage of a collection program is slowest, what each stage costs, and which fixes work. The numbers come from DexSet’s own egocentric and teleoperation pipelines, where we track cost and cycle time per stage across every program.

Key Takeaways – The compliance bottleneck has three chokepoints: consent throughput at recruiting, redaction cost and QA in post, and legal review latency at contracting. – In our pipelines, PII redaction adds $3 to $10 per collected hour, a 10 to 25% overhead on typical teleop collection costs of $28 to $60 per hour. – Consent cannot be retrofitted. Data collected without proper consent is not slow to fix; it is frequently unusable. – Ego4D and Project Aria proved consented egocentric collection scales to thousands of hours, so “compliance kills scale” is an execution complaint, not a law of nature. – The fix is pipelining: run consent, capture, and redaction as concurrent stages with evidence generated inline, not as sequential projects.

What Makes Compliance a Bottleneck Rather Than a Cost?

A bottleneck is the stage of a pipeline that sets the throughput of everything behind it, and in physical AI data programs that stage is almost never the camera. A modern capture rig produces data far faster than the surrounding legal and privacy machinery can clear it. When teams say collection is slow, stage-level tracking usually shows capture idle and one of three other stages saturated.

We see the same three chokepoints across programs. Consent throughput: how many screened, consented, scheduled participants you can put in front of a rig per week. Redaction throughput: how many captured hours per week clear automated PII detection plus human QA. Legal latency: how long consent templates, jurisdiction reviews, and buyer diligence take before anything ships. Each behaves differently, so each needs a different fix.

Chokepoint 1: Consent Throughput at Recruiting

Consent throughput is the rate at which you can convert interested humans into participants with signed, modality-specific, jurisdiction-correct consent on file. It is the least automatable stage. A person has to read a document that names video, audio, hand pose, and gait capture; understand that the data trains commercial AI models; and sign a release that meets the local standard, including BIPA’s written release requirement where Illinois is in scope.

Teams underestimate this because a photography-style release takes five minutes. A defensible AI training consent takes longer per participant and much longer per template, and every jurisdiction you add multiplies the template work. GDPR wants explicit consent for anything approaching Article 9 biometric territory. BIPA wants a written release before collection with statutory damages behind it. The throughput fix is unglamorous: standing consent infrastructure. Pre-approved templates per jurisdiction, digital signing, a participant pool consented once for a program rather than per session, and versioned re-consent when scope changes. Programs with standing infrastructure onboard in days; programs drafting consent per project lose weeks before the first frame.

Chokepoint 2: Redaction Cost and QA in Post

Redaction throughput is the rate at which captured hours clear face, license plate, and screen removal with enough QA that you would defend the result in an audit. Automated detection is good now (Meta’s open-sourced EgoBlur models handle faces and plates well), but automation without sampled human QA is how misses ship: reflections in oven doors, a phone screen picked up mid-task, a name tag no detector was trained for.

In our pipelines the redaction pass adds $3 to $10 per collected hour, with controlled indoor scenes at the low end and bystander-dense scenes at the top. We sample 10 to 20% of episodes for human QA, and 100% for anything captured in public. The cost is manageable; the throughput trap is treating redaction as a batch project after collection ends. Run it inline, per session, and misses feed back into scene sweeps while the program is still running.

Chokepoint 3: Legal Review Latency

Legal latency is the calendar time consumed by consent template approval, jurisdiction mapping, and buyer due-diligence, and it is the chokepoint that kills deals rather than throughput. A buyer’s counsel who cannot verify consent provenance does not negotiate; they walk. We have watched consent documentation decide vendor selection between technically comparable datasets, which is why we treat the audit trail (every episode linked to a consent ID, sweep log, and QA result) as a sales asset, not overhead.

Where the Time and Money Actually Go

The distribution of cost and cycle time across stages is more useful than any single number, so here is the shape of a typical consented egocentric or teleop program in our experience.

StageTypical cost shareCycle-time behaviorRetrofit possible?
Recruiting + consent$1 to $3 per hour collectedWeeks up front; days with standing infrastructureRarely; participants unreachable
Capture (rig + operator)$28 to $60 per teleop hourScales with rigs; almost never the constraintN/A
Scene sweep for bystandersMinutes per sessionInline if scheduled; blocks session if notNo
PII redaction + QA$3 to $10 per hourInline if pipelined; months if batchedYes, at 2 to 3x cost
Audit trail + dataset cardMarginal if built inNear zero inlineWeeks of reconstruction
Legal review + buyer diligenceFixed per programHours with evidence; weeks withoutSometimes deal-ending

Read the last column carefully. Capture capacity recovers from any mistake; consent does not. That asymmetry is the entire argument for consent-first design.

The Bottleneck Is Real, But It Is Not a Law of Nature

The claim that consented collection cannot scale is contradicted by the largest egocentric datasets in existence. Ego4D reached roughly 3,670 hours across 900-plus camera wearers under institution-level ethics review and consent protocols, with face and PII blurring applied where required. Project Aria ran wearable capture with trained wearers, visible recording indicators, and automated bystander anonymization. Neither program treated privacy as a tax on scale; both engineered it as a pipeline stage. Commercial programs have fewer excuses, not more, because we control recruiting, locations, and tooling in ways academic consortia often could not.

The full regulatory map, the six-link consent chain checklist, and the de-identification cost table live in our pillar guide: The Complete Guide to Privacy, Consent & Compliance in Human Data Collection. If you are budgeting a program this quarter, start there.

This article is informational, not legal advice. Consult counsel for decisions about your own data program.

Next Step

Budgeting a collection program? Read the complete privacy, consent and compliance guide, or download the RFP Scorecard to pressure-test your current vendor’s consent chain.

Frequently Asked Questions

Why is human data collection the main bottleneck in physical AI?

Robot models need newly created human demonstration data rather than existing web data, and every created hour requires consent, bystander handling, and PII redaction before it is usable. Those stages, not capture hardware, set program throughput.

In DexSet’s pipelines, consent workflow adds roughly $1 to $3 and PII redaction $3 to $10 per collected hour, together a 15 to 30% overhead on typical teleoperation costs of $28 to $60 per hour.

Usually not. Participants become unreachable, and data collected without proper consent for AI training often cannot be re-licensed, which is why consent-first design beats retrofit in every case we have reviewed.

Build standing consent infrastructure: pre-approved jurisdiction-specific templates, digital signing, a consented participant pool, and versioned re-consent, then run redaction inline with capture rather than as a batch afterward.

Yes. Ego4D collected roughly 3,670 hours under institutional consent and de-identification protocols, and Project Aria ran wearable capture with bystander anonymization built in, both documented publicly.

Enoch Pakanati
Written by

Enoch Pakanati

Enoch Pakanati is the strategic architect behind DexSet’s mission to become the undisputed market leader in robotics training data. He oversees the company’s growth strategy, focusing on capturing dominant market share across all data modalities required for modern robotics, including egocentric capture, teleoperation, and simulation-to-real data pipelines.

At DexSet, Enoch is responsible for transforming the company’s deep technical capabilities into a market-leading brand that foundation model labs and robotics OEMs trust implicitly. He focuses on scaling DexSet’s global footprint and ensuring the company stays ahead of the industry’s rapidly evolving data needs. His leadership is centered on one objective: making DexSet the singular, global standard for the data that powers the robotics revolution.