Most published manipulation data involves rigid objects, and for good reason. A mug has six degrees of freedom. Describe its position and orientation and you have described its state completely. You can reset it to the same spot a thousand times, verify the reset with one overhead camera, and annotate a grasp by pointing at a fixed geometric feature.
A shirt obeys none of this. It has, for practical purposes, an infinite configuration space. Drop the same shirt on the same table twice and you will get two crumple patterns that have never existed before and will never exist again. There is no compact state vector to write down, no canonical pose to reset to, and no fixed feature to anchor a label. Everything that makes rigid-object capture tractable is missing, which is exactly why deformable data commands the price it does.
Self-occlusion is the default condition
Cloth hides itself. The fold you care about happens under a second layer of fabric; the corner the demonstrator is reaching for is tucked inside a crumple; the hand performing the pinch disappears behind the very material it is manipulating. On rigid objects, occlusion is an event you engineer around. On garments it is the permanent operating condition.
This is why single-camera cloth footage is close to worthless for training and why our garment stations are the most heavily instrumented on the floor: overhead multi-view coverage to preserve whatever full-scene truth is recoverable, plus egocentric head and wrist streams to record what the demonstrator actually saw and touched. When labs ask why deformable capture cannot simply be filmed with a phone on a tripod, self-occlusion is most of the answer.
Every touch destroys the state
With rigid objects, a failed grasp usually leaves the scene as it was. With cloth, the act of touching changes the configuration: lift a sleeve and the whole garment re-drapes. There is no undo. This has a practical consequence that surprises people outside the trade: you cannot re-run a cloth episode. Each demonstration is unrepeatable, so protocol quality has to be right the first time, every time.
Because exact resets are impossible, we make the distribution reproducible instead. Our garment protocols script the randomisation itself: drops from a fixed height, defined shake counts, standardised flattening passes for canonical starts. Two sessions a month apart cannot produce identical crumples, but they can produce statistically comparable ones, and that is what a training set actually requires.
Simulation has not closed this gap
For rigid objects, simulation now carries a meaningful share of the training load, and sim-to-real transfer is a routine part of serious pipelines. Cloth is where that strategy thins out. Simulating fabric well means modelling bending stiffness, shear, friction against itself, and contact dynamics through crumpled geometry, and doing it fast enough to generate data at scale. Current simulators approximate this, and the approximation shows: policies trained on simulated cloth meet real fabric and discover that real fabric does not behave.
The industry-wide reading, which matches our experience, is that real capture still carries the load for deformables. That scarcity is structural, not temporary hype, and it is the second reason the premium exists.
Success is a judgment, not a coordinate
A rigid pick either ended with the object in the bin or it did not. A folded shirt is a perceptual outcome: are the edges aligned within tolerance, is the collar square, would a person accept this fold? We grade garment episodes on defined rubrics rather than binary flags, and we keep labelled failures and recoveries in the corpus deliberately, because a policy that has never seen a botched fold corrected has not seen the task. Our note on how capture pipelines actually run covers where this grading sits in the QC chain.
Why labs pay the premium
Scarcity meets demand. The public corpora are overwhelmingly rigid pick-and-place, so deformable coverage is thin exactly where general-purpose policies visibly fail today. Meanwhile the commercial pull is strong: laundry handling, garment folding, bed-making, dressing assistance, and textile handling in industry are all high-value deployment targets, and all of them are gated on cloth competence. When a capability is both commercially urgent and absent from free data, purpose-built capture is the only supply, and it is priced like the specialised work it is.
There is also a compounding factor: deformable work overlaps heavily with fine motor task families. Buttoning a cuff is simultaneously a cloth problem and a sub-centimetre alignment problem. Data that covers both at once is the rarest of all.
What disciplined garment capture looks like
For teams evaluating vendors, these are the practices we consider non-negotiable in our own studio:
- A structured garment library. Items varied deliberately across fabric weight, stiffness, colour, pattern, and size, with each item carrying an identifier that appears in every episode's metadata.
- Scripted initial-state procedures. Randomisation by protocol, so the start-state distribution is reproducible across sessions and operators.
- Multi-view plus egocentric coverage. Enough geometry to survive self-occlusion, time-aligned to one clock.
- Grasp-point annotation on the fabric itself. Labelling where on the garment the demonstrator took purchase, not just when.
- Phase labels anchored to observable events. Flatten, align, first fold, second fold, each boundary tied to something visible in frame.
- Failure and recovery retained. Graded outcomes, with imperfect episodes labelled rather than discarded.
None of these practices is glamorous. Together they are why two datasets of the same nominal size can differ several-fold in training value.