7 Signs Your Training Data Is the Real Bottleneck
Key takeaways
- If failures cluster, the demo-to-field gap is wide, and you can’t state what your dataset covers, data is almost certainly your bottleneck.
- The good news: a data bottleneck is the most fixable one — you can target exactly the variation you’re missing.
Teams reach for a bigger model when the real problem is upstream. Here are seven signs your training data is the bottleneck — the same signals we look for when we diagnose a stalled policy. For the full method, see how to tell if data is your robot’s bottleneck.
1. Failures cluster on specific conditions
Transparent objects, cluttered shelves, low light, the last step of a sequence — when failures concentrate on particular variations, that’s a coverage gap, and coverage gaps are data problems, not capacity problems.
2. It's great in the demo and bad in the field
The demo distribution is narrow; the field is the long tail. A policy that shines in controlled conditions and breaks in deployment is telling you the training data never saw reality. This gap is the structural bottleneck the field keeps rediscovering.
3. There's a big gap between training and held-out performance
Strong on training conditions, weak on a properly separated held-out set means the model memorized a distribution it can’t leave — a generalization problem rooted in data coverage.
4. Performance decays over time
A policy that was fine at launch and degrades over months is suffering distribution shift against a frozen dataset. Without fresh data, models regress. The fix is a data flywheel.
5. Adding more data isn't helping
If throwing volume at the problem does nothing, you’re adding more of what you already have. Coverage — not volume — is the lever.
6. You can't say what your dataset actually covers
If no one can list the variation axes your data spans, you can’t know what’s missing — and what’s missing is what breaks the robot. This is exactly what coverage scoring fixes.
7. Your eval set is suspiciously easy
If your benchmark numbers look great but the robot doesn’t, your evaluation set probably leaked from training. Real eval data is captured in separate sessions with documented provenance.
Next Step
Every sign above points to the same move: measure coverage, capture the exact variation you’re missing, and keep the loop running. Bigger models won’t save a thin dataset.
Sainath Gupta
Sainath Gupta is the visionary leader of DexSet, a company dedicated to building the foundational data layer for the robotics revolution. Sainath recognized early on that the primary bottleneck for physical AI wasn't hardware or model architecture, but the lack of high-quality, scalable training data.
At DexSet, he has pioneered a "manufacturing-first" approach to data collection. This includes the development of transparent cost-per-hour benchmarks and the scaling of global networks for egocentric and teleoperation capture. Sainath is a vocal advocate for the industry’s shift toward treating human experience as the primary pretraining substrate for robots. His leadership at DexSet is focused on one goal: providing the millions of hours of high-fidelity data required to bring humanoid robots out of the lab and into the real world.