World Motion Models

Flexible Sequence Modeling of SE(3) Trajectories

Jiahui Lei1, Qianqian Wang2, Trevor Darrell1, Angjoo Kanazawa1
1UC Berkeley  ·  2Harvard University

NeurIPS 2026 (Spotlight)

World Motion Model
entitiestime
Policy and Model
Predictive Control
3D Scene
Future Prediction
Whole-body Object
Interaction Generation
Human-Humanoid
Retargeting
Hand Object
Interaction Generation
Manipulator
Motion Planning
Swipe to explore all six applications

A Unified Language for Everything in 4D

A set of SE(3) trajectories approximates the 4D world.

Much of what we experience in the world can be approximated by a single SE(3) trajectory or a composition of several trajectories. For example, in the scene above, a single trajectory captures the motion of the camera or rigid cube. Combining trajectories lets us describe articulated furniture or a robot arm, or use skinning to represent the linked motion of a human, hand, or animal. Even general non-rigid, unstructured objects, such as the deforming bear and towel, fit this picture: a small set of SE(3) basis motions can be blended to approximate continuous deformation (DynMF, MoSca, SoM).

This language unifies world states and robot actions.

State and action

Once the local PD gains of the motor controller are fixed, a robot action collapses to a desired SE(3) pose for every link — the forward kinematics of the joint-space PD target (gold ghost). The realized state (blue gripper) is the SE(3) pose each link actually reaches. State and action therefore share the same SE(3) language.

We encode the where, when, and what of each world element as tokens, organize them into a table, and learn a sequence model over this representation.

A pose token at every step

Token

T: 6D pose
Where
t: timestep
When
F: feature
What
Robot
action
Robot
state
Scene
object
Human
joint

Each token is an SE(3) pose paired with a per-trajectory feature, encoding the where (T ∈ SE(3)) and the what (F ∈ ℝD) at each time step. The grid above represents a dynamic 3D scene as an M × N token tensor: rows index the entity and columns index the time step. Time indices are either consecutive within the causal window, or disjoint virtual indices serving as non-causal context (goal frames, retargeting exemplars).

Flexible Sequence Modeling

WMM unifies flexible sequence modeling in a single diffusion model, conditioning on any partially known context.

Training

Per-token noise level corrupted tokens
World
Motion
Model

During training, we corrupt each token independently with a per-token noise level; WMM then predicts the clean tokens.

Inference

    At inference, one trained WMM can condition on any partially known context in the token table and generate the unknown tokens. By changing which tokens are observed, the same model flexibly supports all the modes below.

  • (i) Future Prediction P(s′ | s) / Policy P(a | s)

    Observe history, predict future tokens: state tokens forecast the scene; action tokens make the same weights serve as a policy.

  • (ii) Action-Conditioned Prediction P(s′ | s, a)

    Keep actions observed and predict future states: WMM becomes a learned forward dynamics simulator, or world model, that supports planning and model-predictive control (MPC).

  • (iii) Motion Planning / Infilling

    Observe sparse keyframes and generate the missing motion between them for planning or interpolation.

  • (iv) Retargeting, Inverse Dynamics P(a | s′, s), etc.

    Observe a partial entity set and generate the rest, enabling retargeting, partner prediction, or inverse dynamics.

Network Architecture

Our transformer architecture enhances conditioning on flexible per-token noise levels and scales to thousands of tokens through factored attention and registers.

Applications

One general architecture, trained the same way, handles all of the tasks below.

Learned Retargeting — Human–Humanoid Joint Modeling

Retargeting inference mask
known generated

WMM can retarget human motion to humanoid robots (and vice versa). It models the joint distribution of the human motion and its physically plausible, retargeted humanoid motion. Most importantly, we do not have to tell WMM which human joints correspond to which robot joints: a few paired context exemplars suffice, so the approach is morphology-free. At inference, WMM generates the humanoid motion from these exemplars and the human motion.

Two human-to-Unitree-G1 exemplar pairs in a 2x2 grid: translucent SMPL humans paired with their retargeted G1 robots, each drawn as SE(3) coordinate-frame triads
Input 1 · context exemplars
Observed source human motion drawn as SE(3) coordinate-frame triads along trajectory curves
Input 2 · known human motion
Generated robot link SE(3) trajectories drawn as coordinate-frame triads on the humanoid
Output · generated robot‑link SE(3) trajectories
Inference speed FPS, higher better
GMR
35.0
PHUMA
2.3
Ours
87.0
Tracking success rate higher better
GMR
0.79
PHUMA
0.83
Ours
0.87

Faster and safer retargeting: compared with optimization-based methods — GMR is real-time but less accurate and unsafe; PHUMA is more physically plausible but slow and can fail — our data-driven prior, trained on a curated retargeting dataset, is both faster and safer. We score quality by feeding each kinematic trajectory to a frozen low-level controller (TWIST) and rolling it out in physics simulation.

motion example 1 / 2
GMR PHUMA Ours
Input human

WMM can be used to retarget noisy motion reconstructed from videos:

video example 1 / 2
Mono Real Video
human · Ours · physics sim

Trained on multiple humanoids: swap the context exemplar at inference — human+G1 or human+H1 — and the same model retargets different humanoids, each with a different number of links and morphology.

Unitree G1
Unitree H1

TACO — Hand–Object Joint Modeling

WMM also models hand–object interactions. Using the same joint-modeling formulation as for retargeting, we train one model on TACO and run it in the different inference modes below:

Future Prediction
known generated

green first frame is the known input — everything else is generated

Object to Hand
known generated

green object is the known input — everything else is generated

Hand to Object
known generated

green hand is the known input — everything else is generated

Motion Planning on Many Morphologies

Motion‑planning inference mask
known generated

WMM can plan motion for robot manipulators. It plans collision‑free motion by modeling the motion of each link of the robot and the obstacle. Given the first and last robot states together with the obstacle’s motion in the middle, WMM generates the intervening robot motion.

Known starting arm pose, shown in green, beside the obstacle Known ending arm pose, shown in green, beside the obstacle
Input: starting and ending poses, and obstacle motion
Output: planned motion of each link

One model, many morphologies: a single WMM models robot arms with different morphologies.

Panda
UR3
UR5
UR10
WidowX
Kinova
iiwa14
iiwa7

LangTable — WMM as Policy and Forward Dynamics for Model Predictive Control (MPC)

Policy inference mask
known generated

WMM can act directly as a robot policy. It predicts actions alongside future states. LangTable is a language-conditioned tabletop pushing benchmark. WMM models the SE(3) state of every block together with the end-effector — its state (green) and action (red) — and pins one context time step to the successful ending frame. At inference, WMM observes only a short history and generates the rest; we read the policy directly from the predicted action row.

Input · short history (T=5)
Imagined success goal state
Gen 1 (context) · success goal
Gen 2 · future states + actions

Policy rollout results — please see our paper for quantitative evaluation.

Forward Dynamics mask
known generated
Box to Absolute Location
Plain policy MPC
success rate (%) ↑
76
80
+4 pp
mean steps ↓
92.7
87.2
−5.5
WMM can also serve as a world model for planning. Given the actions as well, the same WMM runs as a forward dynamics simulator and predicts the resulting future states (mask, right). This unlocks model-predictive control: the model first imagines a successful goal (below, left), then scores the action chunk’s final frame against it with a value function based on their difference and optimizes the chunk with both sampling-based and differentiable gradient-based MPC. This enables faster rollouts at a higher success rate versus the plain policy (note the rapid push of the cube in the video below).

Task: “push the blue cube to the top right”

Imagined Goal: the blue cube at the top right
Value Function
(Distance)
Imagined Goal Predicted Chunk Last Frame Optimized Action Rollout
(note the peak push)

OMOMO — Full-Body Human–Object Joint Modeling

WMM can also model humans interacting with objects. Here we demonstrate human–object interaction motion infilling. Given the first frame of the human and object and several waypoints of the object’s location, WMM completes the object’s motion and generates the interacting human motion.

Comparison with CHOIS. On the OMOMO dataset, our generation looks more natural than that of CHOIS (ECCV 2024) — even though WMM is not tailored or optimized specifically for this task. Please see the paper for the quantitative comparison.

Motion diversity — one object trajectory, six random seeds. WMM produces varied yet plausible ways to perform the same object‑conditioned interaction.

TraceGen — Dense 3D Scene Future Prediction

First‑frame inference mask
first frame (t = 0) generated, 32 steps

WMM can predict the 3D future of general scenes. It handles dense prediction, querying hundreds of trajectories at once. This formulation subsumes existing scene-future-prediction works such as MoMaps, TraceGen, and PointWorld. At inference, the input is the first-frame RGB‑D image; the prediction is the future of the tracks at dense query positions.

TraceGen input: first-frame RGB-D — 2D camera view with dense query points beside the 3D point cloud

Known input: first-frame RGB‑D

Generated output: future 3D tracks

Results — each scene shows 2D camera tracks (top) and 3D point‑cloud tracks (bottom); the two columns compare the TraceGen baseline and Ours.

3D‑trajectory error vs. the TraceGen baseline
TraceGen Ours
Droid ↓ lower is better (×100)
0.206
0.151
MSE
27% improvement
1.289
1.129
MAE
12% improvement
0.285
0.200
Endp.
30% improvement
EpicKitchen ↓ lower is better (×100)
0.445
0.322
MSE
28% improvement
2.721
2.130
MAE
22% improvement
0.791
0.556
Endp.
30% improvement

Information

BibTeX
@inproceedings{lei2026wmm,
  title     = {World Motion Models: Flexible Sequence Modeling of SE(3) Trajectories},
  author    = {Lei, Jiahui and Wang, Qianqian and Darrell, Trevor and Kanazawa, Angjoo},
  booktitle = {Advances in Neural Information Processing Systems},
  year      = {2026}
}

Acknowledgements

This project was funded in part by Meta BAIR partners, the BAIR Humanoid Intelligence Center (BAIR HIC), and NSF CAREER (No. 2442491).