How to Sell Robot Training Data
How to sell robot training data that buyers will actually license: schema, consent basis, provenance, and the evaluation slice that turns hours into an asset.
A capture shop with 800 hours of pick-and-place footage on a drive emails three model teams. Two never reply. The third asks one question that ends the pitch: how was consent obtained, and can we test a sample before we commit? The footage was real. The recording was clean. But it was not yet something a buyer could license, and the seller had confused the two.
That confusion is the most expensive mistake a data seller makes. Hours on a disk are inventory. A dataset a buyer will pay for is a product, and the gap between them is filled with unglamorous work most capture operations underestimate. This is a guide to closing that gap: what to add to raw footage, how to get discovered without giving the data away, and which pricing and rights choices you only get to make once.
From hours to an asset: what a buyer is really paying for
A buyer does not want your video. They want a predictable improvement to a policy, delivered in a format their pipeline already reads, with a title clean enough to survive their lawyer. Raw hours deliver none of that on their own. The work that turns capture into a licensable asset is specific, and it is checkable.
| Aspect | Raw hours on a drive | What the seller adds |
|---|---|---|
| Format | Mixed codecs, one folder per rig | One documented schema a loader accepts |
| Segmentation | One long unbroken recording | Episodes cut and labeled by skill |
| Provenance | We filmed it, trust us | Per-session consent record and a stated jurisdiction |
| Verification | No way to check quality first | A random eval slice a buyer can test |
| Integrity | Loose files, no checks | A manifest whose hashes match the bytes shipped |
Two open efforts show why format is not optional. Aggregations such as Open X-Embodiment unified dozens of datasets into one action format so a model could train across them, and tooling like Hugging Face LeRobot made publishing to a common schema cheap. If your trajectories arrive in a shape a buyer has to reverse-engineer, the integration cost swamps the price of the data, and the deal dies on effort alone.
Labeling is where most of the value gets added, and where most sellers stop too early. A buyer can use raw video for pretraining, but a policy that has to act needs episodes segmented by skill, aligned across sensors, and tagged with the moment contact begins. That annotation is domain work, not a checkbox, and it is the difference between a corpus a team can drop into a pipeline and one they have to rebuild before it earns its price.
The eval slice: the one thing that lets a stranger price your data
No serious buyer wires money for data they cannot inspect. But you cannot hand over the whole dataset for evaluation, because then there is nothing left to sell. The resolution the market is converging on is the evaluation slice: a small, random, representative sample a buyer can train or probe against under controlled terms before committing to the rest.
The word random matters. A curated highlight reel proves nothing except that you can pick your best minutes. A slice drawn at random from the full corpus is a statistical promise about the whole. Pair it with a documented capture protocol, and a buyer who has never met you can still form a price. Datasets like DROID earned trust partly by publishing their protocol openly, so anyone could judge what the data was and was not.
Getting discovered without giving the data away
Sellers worry that listing their data means exposing it. In business-to-business data trade, that worry is handled by masked discovery, and it is completely normal. A buyer browses a structured dataset card, not the files: hours, embodiment, sensor modalities, tasks, capture rate, jurisdiction, rights. Identity and raw bytes stay behind a gate until scoping is serious and terms are moving.
This is not evasion. It is how sensitive assets have always changed hands. The card carries enough for a buyer to decide whether the data could fit; the eval slice, released under terms, confirms it. Anonymized discovery also protects the people in your footage, which matters more the moment human demonstration is involved.
Pricing and rights: the choices you make once
Price follows structure, not size. The deployment curve tracked by the International Federation of Robotics keeps demand high, so a seller who has solved trust rarely lacks buyers. What varies is the rights bundle, and each choice is close to irreversible.
Exclusive or non-exclusive. Sell the same dataset to ten buyers and it is cheap and repeatable; sell it once, exclusively, and it commands a premium but cannot be resold. Commercial or research-only. Whether the buyer may train a shipping model, or only publish a paper, is a different asset at a different price. Derivative rights. A model trained on your data is a derivative; whether that output is fenced in or set free is a term you write, not an afterthought.
One constraint sits above all of these. You can only license rights you actually hold. If the consent you gathered did not cover commercial training, no clause can invent it later. The EU Data Act is widening who may share machine-generated data, but it does not paper over a missing consent basis for the personal data inside a human demonstration. Rights you cannot document are rights you cannot sell.
Raw hours are inventory. A dataset is inventory plus a schema, a consent record, and a sample a stranger can test. The last three are the product.
The practical sequence is the same for a capture shop, a lab, or a factory sitting on old recordings. Decide the schema before you record, not after. Reserve a random slice you will never sell, so buyers can always test. Gather consent that covers the use you intend to license. Do those three things and your hours become an asset a stranger can price. Skip them and you are left doing what the shop in the opening did: emailing footage no one can buy.