World-model data glossary

World action model

Definition

World action model is an emerging term for a system that makes environment dynamics and action part of the same learning problem. Depending on the implementation, actions can condition a predicted future, be predicted jointly with video or state, or act as policy outputs. The phrase does not define one standard architecture, control level, or dataset schema, so a model description still needs to state what it observes, what it predicts, and how action enters the interface.

Why it matters

The category is useful only when action is technically inspectable. Buyers should ask whether controls were recorded or inferred, which embodiment and coordinate system they belong to, how timestamps align, whether commanded and executed actions differ, and which outcomes are observed. A passive video collection does not become action data through naming. The evaluation should test intervention sensitivity, offline action prediction, closed-loop behavior where available, and transfer to held-out tasks or embodiments without collapsing those results into one marketing metric.

Example

A system receives a task instruction, recent observations, proprioception, and an action history. It predicts an action chunk together with the future scene dynamics those actions should produce. A robot policy can then execute the action output without generating video at deployment, while the joint future-prediction objective still shapes the learned representation during training. The exact evidence required depends on the claimed interface and evaluation setting.

Primary sources

Continue

Apply this definition to dataset scope, fields, rights, and evaluation.

World action model data