The Humanoid Data Company: Why Proprietary Data Wins
Hardware is converging, so the humanoid moat moved to proprietary robot data. Why first-person demonstration data is the asset that compounds.
Walk onto the floor of almost any humanoid startup in 2026 and you find the thing the pitch deck leaves out: rows of people. Some wear VR headsets and drive a pair of robot arms by hand, folding the same towel for the two-hundredth time. Others move through a motion-capture volume in instrumented suits. The robot is what the investors came to watch. The humans generating trajectories are the actual product line.
For a decade the hard part of a humanoid was the body: actuators that survive a fall, hands with real fingers, a battery that lasts a shift. That problem is not finished, but it is converging. Figure, 1X, Agility Robotics, and Boston Dynamics now ship platforms whose spec sheets rhyme. When hardware converges, the quiet question inside every serious team shifts from "can we build the robot" to "what will it learn from, and who else has that data."
The answer is not just any data. It is proprietary records of embodied behavior: how a hand approaches a mug, how grip force changes the instant a lid resists, where a person looks half a second before they reach. That asset compounds. Model architectures leak in months. A behavior corpus you collected and no competitor holds does not.
The hardware stopped being the hard part
Look at what shipped in the last two years and the pattern is hard to miss: multiple vendors, comparable degrees of freedom, comparable hand designs, similar battery envelopes. The mechanical frontier still matters, but it no longer separates the leaders from the pack the way it did in 2021. What separates them is the policy running on top, and the policy is downstream of what it was trained on.
This is the lesson the language-model era already taught, transposed onto atoms. Transformers were a public paper within weeks. What kept the frontier labs ahead was corpus, cleaning, and feedback data nobody could copy. Robotics is now living the same story with one brutal difference: there is no web-scale corpus of physical interaction lying around to scrape. Every useful trajectory has to be produced.
What the open datasets do and do not give you
The field's answer to scarcity has been to pool. Open X-Embodiment gathered more than a million real robot trajectories across 22 embodiments from 21 institutions, and showed that a policy trained on the mixture beats one trained on any single lab's slice. DROID added roughly 76,000 teleoperated trajectories across 564 scenes, deliberately diverse in setting. These are gifts to the field. They are also, by construction, shared: your competitor downloaded the same files.
The table below sketches the tradeoffs. Read the last column first.
| Dataset | Approximate scale | Source | What it does not give you |
|---|---|---|---|
| Open X-Embodiment | 1M+ trajectories, 22 embodiments | Pooled robot teleoperation | Your task, your hand, your edge cases |
| DROID | ~76,000 trajectories, 564 scenes | Franka teleoperation | Non-tabletop, contact-rich, bimanual work |
| RH20T | On the order of 100,000 episodes | Contact-rich teleoperation | Human-native motion and first-person gaze |
| Ego4D | ~3,600 hours of video | First-person human, in the wild | Actions, forces, robot-usable labels |
| Ego-Exo4D | ~1,000+ paired hours | First-person plus third-person human | Contact forces and a shared action space |
Two gaps recur. First, almost all robot data is teleoperated on tabletops, which under-samples the contact-rich, whole-body, two-handed work humanoids are sold to do. Second, the large first-person human datasets, Ego-Exo4D among them, are rich in perception but thin on the action and force labels a policy needs to imitate. Bridging the two is unsolved, and it is exactly where a proprietary corpus earns its keep.
Why proprietary data compounds when models do not
Deployment is where the flywheel turns. A robot in a real environment generates observations no dataset anticipated: the failure at 2am, the glare on a steel counter, the object left where it should not be. Physical Intelligence has argued for generalist policies trained across many tasks and platforms; Toyota Research Institute frames its Large Behavior Models the same way, as behavior learned at scale rather than scripted. Both bets assume a data engine feeding them. NVIDIA's Isaac GR00T is explicit that a foundation model is only as good as the mixture of real, human, and synthetic motion behind it.
Here is the compounding: a better policy earns more deployments, more deployments produce more and rarer data, that data trains a policy competitors cannot match because they never saw those situations. The model weights are a snapshot. The pipeline that keeps refilling them is the moat.
The model you ship is a snapshot. The pipeline that keeps refilling it is the asset. Weights can be distilled; a data flywheel has to be lived.
The cost structure nobody puts on a slide
Teleoperation, the default source of high-quality action data, is expensive in a way founders rarely disclose. Every trajectory is a human operating a robot in something close to real time. Throughput is bounded by human hours, and the good data, the recoveries and the corner cases, is the slowest to collect because it is the rarest to encounter. Simulation helps for some skills and lies convincingly for others: contact and deformable objects remain where sim-to-real still bites.
First-person human capture flips part of the equation. A person wearing sensors can demonstrate a task at human speed, with human dexterity, far faster than they could puppet a robot. The catch is the embodiment gap: a human hand is not a robot gripper, and retargeting that motion onto a different body is its own research problem. The economics are attractive; the transfer is the work. Anyone claiming both are solved is selling something.
Data as a strategic asset: rights, provenance, jurisdiction
Once behavior data is the balance-sheet item, its legal shape starts to matter. Who owns a trajectory captured from a contracted demonstrator? Under what license can it be resold, pooled, or trained on? The EU Data Act and the EU AI Act push toward documented provenance and clear data-governance duties, and buyers of robot foundation models will increasingly ask where the training data came from and whether the rights are clean. A corpus with murky provenance is a liability wearing the costume of an asset.
This is the unglamorous half of "data company." It is not only collecting: it is consent, labeling standards, retention, auditability, and the ability to prove a chain of custody when a partner's compliance team asks. Teams that treat this as paperwork will discover, late, that it was product.
So when a humanoid company tells you its edge is the robot, ask a sharper question: show me the data engine. Ask what it captures that the open datasets do not, how fast it refills, who owns the output, and what happens to the corpus when the current model is three versions old. The teams that can answer cleanly are already running as data companies. The rest are buying hardware and hoping. In this market the robot is the thing you can see. The data is the thing that pays.