Embodiment changes the meaning of every action
The same high-level instruction can require different kinematics, sensors, control rates, and safety constraints across a single arm, mobile manipulator, bimanual platform, quadruped, or simulated agent.
A release should identify morphology, joints, end effectors, camera placement, calibration, controller, coordinate frames, payload assumptions, and software or firmware version.
Demonstrations need context and negative evidence
Successful demonstrations teach one route through a task. Failures, recoveries, corrections, aborted attempts, and human interventions reveal decision boundaries and operational limits.
Episode records should preserve who or what controlled the system, how instructions were delivered, whether assistance occurred, why the episode ended, and which success test was applied.
- Teleoperation, autonomous, scripted, or mixed control
- Task instruction, scene, object, and initial-state metadata
- Raw sensors, calibrated observations, and validity masks
- Success, failure, correction, intervention, and safety events
A unified format should not erase source differences
Cross-dataset loaders make large mixtures practical. They should not discard original action semantics, sampling rates, licenses, collection methods, or quality decisions.
Normalize through explicit adapters and retain source-level provenance. Buyers and researchers need to know which transformations enable comparison and which differences remain material.