Skip to main content

Dexset

The Complete Guide to Privacy, Consent & Compliance in Human Data Collection (2026)

Sainath Gupta

This guide is for informational purposes only and is not legal advice. Consult qualified counsel before designing a collection program.

Consent in human data collection is usually defined as a signed release form; that definition hides the real problem. The form covers the person who signed it. It says nothing about the roommate who walks through frame in week three, the voiceprint riding in the comms audio, or the gait signature no blur tool can remove because body movement is the very thing the model is learning. Teams that treat consent as paperwork keep discovering, at due-diligence time, how much of their corpus the paperwork never touched.

The gap exists because physical AI inverted the usual data pipeline. Language models trained on text that was already public. Robot models train on video that has to be created, by real people, in real spaces, often wearing head-mounted cameras that record everything in front of them. There is no scraping shortcut. Someone has to recruit those people, obtain informed consent scoped to every signal captured, sweep the scene for bystanders, and prove all of it later when a buyer, an auditor, or a regulator asks.

That last verb is this guide’s thesis: privacy, consent, and compliance are evidence problems, not paperwork problems. A program is defensible when it can show, for any episode on demand, who consented to what, what was redacted, and under which terms the data shipped. This guide gives you the full map for building that: which laws actually apply to robotics human data (GDPR, CCPA/CPRA, Illinois BIPA, the EU AI Act), what the academic community already established with Ego4D and Project Aria, what de-identification really costs per hour, and a six-link consent chain you can implement this quarter. We collect egocentric, exocentric, teleoperation, mono, and stereo data for a living at DexSet, so the numbers here come from our own pipelines, not from a compliance vendor’s brochure.

TL;DR: Key Takeaways – Human data for robot training routinely touches regulated categories: faces, voices, gait, home interiors, and bystanders. Under GDPR Article 9, biometric data processed to uniquely identify a person is a special category requiring explicit consent or another narrow exception. – Illinois BIPA is the sharpest US risk for biometric collection: it requires a written release before collection and carries statutory damages of $1,000 per negligent violation and $5,000 per intentional or reckless violation, with a private right of action. – The EU AI Act (Regulation (EU) 2024/1689) adds data governance and documentation duties for high-risk systems and transparency duties for general-purpose AI model providers, including training data summaries. – In our pipelines, PII redaction adds roughly $3 to $10 per collected hour depending on scene density and QA depth. Budget for it up front; retrofitting is 2 to 3 times more expensive. – The single most useful artifact you can demand from any data vendor is a consent chain: recruit records, informed consent docs, scene sweep logs, redaction QA logs, license terms in the dataset card, and an audit trail linking every episode to a consent ID.

What Is Privacy, Consent and Compliance in Human Data Collection?

Privacy, consent, and compliance in human data collection is the set of legal, contractual, and operational controls that make it lawful to record identifiable people, use those recordings to train AI models, and prove both facts on demand. The three words carry distinct work. Privacy is about minimizing what you capture and who can be identified in it. Consent is the documented permission of the people you record, scoped to specific uses such as model training and commercial licensing. Compliance is the evidence layer: the policies, logs, and audits that demonstrate you did what the law and your contracts require.

For robotics teams, the practical unit of analysis is the episode: one continuous recording of a human demonstration. Every episode has a camera wearer or operator (a consented participant), possibly bystanders (usually not consented), a location (often a private home or workplace), and multiple identifying signals inside it (face, voice, gait, tattoos, screens, mail on the counter). A compliant program accounts for all of them, per episode, at collection time.

Privacy: Minimize What You Capture

Privacy in a capture program means recording only what the model needs and stripping identity signals from everything else. Egocentric rigs are indiscriminate by design; a head-mounted camera records whatever the participant looks at. The privacy work happens in scene control (choosing and sweeping locations), sensor configuration (do you actually need audio?), and post-capture redaction (faces, plates, screens). Data minimization is not just a GDPR principle (Article 5(1)(c)); it also cuts redaction cost, because every identifying signal you avoid capturing is a signal you never pay to remove.

Consent: Documented, Specific, Revocable

Consent for AI training data means a written, plain-language agreement that names the modalities recorded, the uses permitted (including commercial model training and dataset licensing), the retention period, and how the participant can withdraw. Generic photography releases fail this bar. Under GDPR, consent must be freely given, specific, informed, and unambiguous (Article 4(11)), and explicit where special category data like identifying biometrics is involved (Article 9(2)(a)). Under Illinois BIPA, biometric identifiers require a written release before collection, full stop. Withdrawal is the part most teams under-engineer: if a participant revokes, you need to find and delete their episodes, which requires the audit trail described later in this guide.

Compliance: Evidence or It Did Not Happen

Compliance is the ability to produce evidence that your privacy and consent controls actually ran, for every episode, on demand. In practice this means an immutable log linking each episode ID to a consent record ID, a scene sweep record, a redaction QA result, and the license terms it shipped under. Buyers increasingly ask for this during procurement. In the RFPs we responded to over the past year, consent documentation questions appeared in the majority of them, and in several cases the consent chain was the deciding factor between otherwise similar vendors. It is now a differentiator, not a formality.

Why Robotics Human Data Is Different From Web Data

Robotics human data differs from scraped web data because it is created rather than found, biometric-adjacent by nature, and captured in private spaces where bystanders have no relationship with you at all. That combination triggers obligations that web-scale text pipelines never faced.

Three properties matter most. First, the identity signal is dense: a single egocentric episode can contain the participant’s hands, voice, reflection, home interior, and family members. Second, some signals cannot be redacted without destroying training value. You can blur a face; you cannot blur gait, hand morphology, or body kinematics, because those are the very things a manipulation or humanoid model is learning from. Where redaction is impossible, consent has to carry the full weight. Third, collection happens continuously across sessions, which matters legally: the Illinois Supreme Court held in Cothron v. White Castle (2023) that BIPA claims accrue with each scan or collection, not just the first, and a 2024 amendment (SB 2979) was needed to limit damages to a single recovery per person per method. Repeated capture of the same participant is the normal operating mode of a robotics data program, so this line of cases is directly relevant.

Entity relationships worth keeping straight: egocentric capture (Ego4D, Project Aria) and teleoperation (ALOHA-style rigs) both produce human demonstration data; that data feeds imitation learning; imitation learning trains VLA models (OpenVLA, pi0, GR00T-class systems). Compliance obligations attach at the capture step but follow the data through every downstream model.

The Regulatory Map for Robotics Training Data

The regulatory map for robotics human data in 2026 is dominated by four frameworks: GDPR in the EU/EEA, CCPA/CPRA in California, BIPA in Illinois, and the EU AI Act layered on top for AI-specific duties. Most collection programs of any scale touch at least two of them.

FrameworkJurisdictionWhat triggers it for robot dataKey dutiesMaximum exposure
GDPR (Reg. 2016/679)EU/EEA residentsAny identifiable person in the data; identifying biometrics are special category under Art. 9Lawful basis (Art. 6), explicit consent for biometrics, data subject rights (access, erasure, portability), DPIA for high-risk processingUp to EUR 20M or 4% of global annual turnover
CCPA/CPRACalifornia consumers; businesses over thresholdsCollecting personal or sensitive personal information (biometric info is sensitive PI)Notice at collection, rights to know/delete/correct, opt-out of sale or sharing, limit use of sensitive PI$2,500 per violation, $7,500 per intentional violation; private action for breaches only
BIPA (740 ILCS 14)Illinois residentsCollecting biometric identifiers: face geometry, voiceprints, fingerprints, iris scansWritten policy, written release before collection, retention schedule, no profit from biometrics$1,000 per negligent violation, $5,000 per intentional or reckless violation; private right of action
EU AI Act (Reg. 2024/1689)AI systems placed on the EU marketHigh-risk system classification; general-purpose AI model provisionArt. 10 data governance for high-risk (quality, bias, provenance); GPAI training data summaries and documentationUp to EUR 35M or 7% of turnover for prohibited practices; lower tiers for other breaches

GDPR: Lawful Basis and the Biometric Question

GDPR requires a lawful basis under Article 6 for any processing of personal data, and robotics collection programs almost always rely on consent because the alternatives fit poorly. Legitimate interest is hard to defend for commercial biometric-adjacent capture, and contract necessity does not cover bystanders. The sharper question is Article 9: biometric data processed “for the purpose of uniquely identifying a natural person” is special category data. Raw video of a face is not automatically special category, but the moment your pipeline runs face embedding, re-identification, or participant matching, you are plausibly in Article 9 territory and need explicit consent. Our position: draft consent to the explicit standard regardless, because you cannot predict what downstream processing a model lab will run.

Data subject rights are the operational load: access (Art. 15), erasure (Art. 17), and portability (Art. 20) all require you to locate one person’s data inside a dataset of thousands of hours. Without episode-level consent IDs, an erasure request becomes a forensic project.

CCPA/CPRA: Sensitive Personal Information

The CCPA as amended by the CPRA gives California consumers rights over personal information and creates a “sensitive personal information” category that includes biometric information processed to identify a consumer. For data vendors the practical duties are notice at collection, honoring deletion requests, and being careful with the definitions of “sell” and “share” when licensing datasets, because dataset licensing to a model lab can qualify. Contract terms with buyers should pass deletion obligations downstream.

BIPA: The Statute With Teeth

Illinois BIPA is the strictest biometric law in the US because it combines a written consent requirement before collection with a private right of action and statutory damages. Face geometry and voiceprints are both covered identifiers. The math is what gets attention: at $1,000 to $5,000 per violation, a collection program touching a few hundred Illinois residents without written releases carries seven-figure theoretical exposure. Post-2024, damages are capped at a single recovery per person per collection method, but the consent requirement is unchanged. If you collect in or recruit from Illinois, the written release is non-negotiable, and your consent form should be reviewed specifically against BIPA’s language.

EU AI Act: Duties for the Model, Not Just the Data

The EU AI Act regulates AI systems and models rather than data directly, but it reaches back into your dataset through documentation duties. Providers of high-risk AI systems must meet Article 10 data governance requirements covering training data relevance, representativeness, error checking, and provenance. Providers of general-purpose AI models must maintain technical documentation and publish summaries of training content. If your customers are building GPAI or high-risk systems, they will ask you for provenance documentation because they need it for their own filings. A vendor who cannot produce a consent chain makes their customer’s AI Act paperwork impossible.

What the Academic Community Already Proved: Ego4D, EgoExo4D, and Project Aria

The academic precedent for consented egocentric collection at scale is Ego4D, EgoExo4D, and Meta’s Project Aria program, and together they demonstrate that rigorous consent and de-identification are compatible with thousands of hours of useful data. Teams designing commercial programs should study them before inventing anything.

Ego4D collected roughly 3,670 hours of egocentric video from over 900 camera wearers across multiple countries, with each partner institution operating under its own ethics review and consent protocols, and with de-identification steps such as face and PII blurring applied where participants and bystanders had not consented to identifiable release. EgoExo4D extended the model to synchronized egocentric and exocentric capture of skilled activities with consented participants across a multi-institution consortium. Project Aria built privacy into the hardware and workflow: a visible recording indicator, trained wearers, restrictions on where recording happens, and automated anonymization of bystander faces and license plates using models Meta later open-sourced as EgoBlur.

The lesson is not that you should copy any single protocol. It is that the three hard problems (participant consent, bystander protection, and scalable redaction) all have working reference implementations with published documentation. When a vendor tells you consented collection cannot scale, these projects are the counterexample.

The DexSet Consent Chain: Six Links, In Order

The consent chain is our name for the six sequential controls that take a collection program from recruiting to a licensable episode, with evidence generated at every step. Break any link and the episodes downstream of the break are compromised. This is the checklist we run internally and the one we suggest buyers demand from any vendor.

  • Recruit. Screen participants for age, jurisdiction, and language. Disclose compensation, session structure, and the commercial nature of the work before anyone signs anything. Log recruiting source per participant.
  • Informed consent document. Plain-language, modality-specific consent naming exactly what is recorded (video, audio, hand pose, gait, depth), what it will be used for (training commercial AI models, dataset licensing), retention period, and the withdrawal mechanism. Written release language reviewed against BIPA where any Illinois nexus exists; explicit-consent standard for GDPR. One signed record per participant, versioned when terms change.
  • Scene sweep for bystanders. A pre-capture walkthrough of the location: identify household members and passersby, obtain bystander consent or exclude them from frame, remove or mask visible documents, mail, screens, and photos of third parties. Logged with a per-session checklist. Sessions involving minors in frame require guardian consent or do not happen.
  • PII redaction pass. Automated detection of faces, license plates, and screens on every frame, followed by human QA on a sampled basis (we sample 10 to 20% of episodes, 100% for public or semi-public scenes). Misses are corrected and fed back to the detector. Every episode gets a redaction status and QA score.
  • License terms in the dataset card. The dataset card records permitted uses, sublicensing limits, retention obligations, and the consent scope the data was collected under, following the documentation norms of Hugging Face dataset cards. A buyer should be able to read the card and know exactly what they may do with the data.
  • Audit trail. An append-only log linking every episode ID to a consent record ID, scene sweep record, redaction QA result, and delivery manifest. This is what makes GDPR erasure requests, BIPA audits, and buyer due-diligence answerable in hours instead of weeks.

Print this list. It is the whole program in six lines; everything else in this guide is implementation detail.

De-Identification Techniques: What They Cost and Where They Fail

De-identification for robotics data means removing or degrading identity signals in recorded episodes while preserving the physical information a model trains on. No single technique covers everything, and the most important fact in this section is that some signals cannot be redacted at all.

SignalTechniqueAutomation maturityEffort and cost per collected hour (our pipeline)Failure modes
FacesDetection + blur or synthetic face swapHigh (open models like EgoBlur)$2 to $4 including sampled QAProfile views, motion blur, reflections in mirrors and appliances
License platesDetection + blurHigh$0.50 to $1Angled plates, non-standard formats, dashcam-style distance
VoiceMute, pitch-shift, or transcript scrubMedium$1 to $2Names spoken mid-task, voiceprint survives naive pitch shift
Screens and documentsDetection + maskLow to medium$1 to $3Phones picked up mid-episode, whiteboards, mail on counters; highest manual miss rate
Tattoos, badges, name tagsManual review + maskLow$0.50 to $2, scene-dependentOnly caught in QA; invisible to standard detectors
Gait and body kinematicsNone viableN/AN/ACannot redact without destroying training value; must be covered by consent

Summing the table explains the number we quote buyers: in our pipelines, a full PII redaction pass adds roughly $3 to $10 per collected hour. Controlled indoor scenes with no bystanders sit at the bottom of that range. Public or semi-public capture with heavy QA sits at the top. Against a typical teleoperation collection cost of $28 to $60 per hour, redaction is a 10 to 25% overhead, which is why it needs its own line in your budget rather than being absorbed into a vague QA figure.

The gait row deserves emphasis. Body movement is identifying (gait recognition is an active research field) and it is precisely what humanoid and manipulation models learn from. There is no redaction answer. The only honest treatments are consent that explicitly covers body movement data, and contractual limits on re-identification attempts downstream.

The Economics of Compliance

The economics of compliance come down to one comparison: paying for consent and redaction at collection time versus paying to retrofit, relabel, or discard data later. We have seen both paths, and retrofit always loses.

Cost lineConsent-first programRetrofit after collection
Consent workflow (recruiting overhead, documents, signing)$1 to $3 per hour collectedOften impossible; participants unreachable
PII redaction$3 to $10 per hour$6 to $20 per hour (no scene control, denser unknowns)
Audit trailMarginal if built inReconstruction project, weeks of engineering
Data written off as unusableNear zero10 to 40% of affected corpus in cases we have reviewed
Buyer due-diligence responseHoursWeeks, sometimes deal-ending

Round numbers for planning: a consent-first program adds roughly 15 to 30% to raw collection cost. That is real money at thousands of hours. It is also the cheapest insurance available against the two expensive outcomes: a corpus your customers’ lawyers will not touch, and a statutory damages claim in a private-right-of-action state.

Case Study Proof: A Humanoid Foundation Model Team

One engagement pattern shows how the chain performs under real deadlines. A humanoid foundation model team needed several thousand hours of consented egocentric and teleoperation data across household environments in three jurisdictions, on a timeline that left no room for a compliance redo. We ran the six-link chain from day one: jurisdiction-specific consent documents (including BIPA-standard written releases for US collection), scene sweeps at every household session, automated face and plate redaction with 15% human QA sampling, and an episode-to-consent audit trail delivered alongside the data.

The outcomes that mattered to the buyer: their counsel approved the corpus for commercial training on first review, their EU AI Act documentation could cite our provenance records directly, and when one participant later withdrew, the deletion was executed and certified within days because every affected episode was already indexed to that consent ID. Redaction ran at roughly $4 per hour on average across the program, at the low end of our range, because scene sweeps kept bystander density down. Compliance did not slow the program; unplanned compliance is what slows programs.

Buyer Due-Diligence: 10 Questions to Ask Any Data Vendor

Buyer due-diligence for human data means verifying the consent chain exists before you sign, because after delivery your model has already eaten the risk. These are the ten questions we believe every RFP for human robotics data should include. We answer all ten in writing; a vendor who cannot should concern you.

  • Can you produce the signed consent record for any specific episode I point to, within 48 hours?
  • Does your consent document name commercial model training and dataset licensing explicitly?
  • Do you use written releases meeting the Illinois BIPA standard wherever there is an Illinois nexus?
  • What is your bystander policy, and what evidence do you keep that it ran per session?
  • What is redacted (faces, plates, voice, screens), by what method, and what is your human QA sampling rate?
  • How do you handle signals that cannot be redacted, such as gait and body kinematics?
  • If a participant withdraws consent, describe the deletion workflow and its typical turnaround, including my obligations as the buyer.
  • Which jurisdictions did collection occur in, and which legal frameworks did you map each site against?
  • Do your dataset cards state license terms, consent scope, and retention obligations?
  • Will you contractually stand behind the consent chain with indemnification for consent defects?

Get the RFP Scorecard

We turned the ten questions above into a weighted scorecard: the DexSet Consent & Compliance RFP Scorecard, a one-page template that lets a Head of Data compare vendors on consent documentation the same way they compare on price and fps. It includes suggested weightings, red-flag answers, and contract language prompts for indemnification and deletion pass-through. Download it, use it against us too.

Next Step

Ready to see a consent chain in production? Book a demo and we will walk you through a live episode-to-consent audit trail, or download the Consent & Compliance RFP Scorecard and put your current vendors to the test.

Frequently Asked Questions

Is egocentric video biometric data under GDPR?

Raw video containing a face is personal data, but it becomes special category biometric data under GDPR Article 9 when processed for the purpose of uniquely identifying a person, for example through face embeddings or participant re-identification. Because you rarely control downstream processing, collecting to the explicit consent standard is the safer design.

In our pipelines, a full redaction pass (faces, license plates, screens, with sampled human QA) adds roughly $3 to $10 per collected hour. Controlled indoor scenes sit at the low end; public scenes with dense bystanders sit at the high end.

Bystanders are identifiable people, so they are protected under GDPR and similar laws even though they never signed anything. The standard treatments are obtaining bystander consent, excluding them via scene sweeps, or anonymizing them in post, which is the approach Project Aria demonstrated with automated face and license plate blurring.

BIPA requires a written release before collecting biometric identifiers such as face geometry or voiceprints, and it pairs that duty with a private right of action and statutory damages of $1,000 to $5,000 per violation. Individuals can sue directly, which is rare in US privacy law and is why BIPA drives most US biometric litigation.

The AI Act’s duties fall mainly on providers and deployers of AI systems and general-purpose models, not on data vendors as such. It reaches vendors indirectly: customers subject to Article 10 data governance or GPAI documentation duties need provenance and consent evidence from their suppliers to meet their own obligations.

Withdrawal generally obligates you to stop processing and delete the person’s data going forward; whether trained model weights are affected is legally unsettled and one reason deletion workflows and clear retention terms matter. An episode-level audit trail makes the deletion itself fast, and buyer contracts should pass the deletion obligation downstream.

Sainath Gupta
Written by

Sainath Gupta

Sainath Gupta is the visionary leader of DexSet, a company dedicated to building the foundational data layer for the robotics revolution. Sainath recognized early on that the primary bottleneck for physical AI wasn't hardware or model architecture, but the lack of high-quality, scalable training data.

At DexSet, he has pioneered a "manufacturing-first" approach to data collection. This includes the development of transparent cost-per-hour benchmarks and the scaling of global networks for egocentric and teleoperation capture. Sainath is a vocal advocate for the industry’s shift toward treating human experience as the primary pretraining substrate for robots. His leadership at DexSet is focused on one goal: providing the millions of hours of high-fidelity data required to bring humanoid robots out of the lab and into the real world.