A set of SE(3) trajectories approximates the 4D world.
drag to orbit · scroll to read on
Much of what we experience in the world can be approximated by a single
SE(3) trajectory or a composition of several trajectories.
For example, in the scene above, a single trajectory captures the motion of the
camera or rigid cube. Combining trajectories lets us describe
articulated furniture or a robot arm, or use skinning to represent the
linked motion of a human, hand, or animal. Even general non-rigid,
unstructured objects, such as the deforming bear and towel,
fit this picture: a small set of SE(3) basis motions can be blended
to approximate continuous deformation
(DynMF,
MoSca,
SoM).
This language unifies world states and robot actions.
State and action
Once the local PD gains of the motor controller are fixed, a robot action collapses
to a desired SE(3) pose for every link — the
forward kinematics of the joint-space PD target
(gold ghost).
The realized state
(blue gripper)
is the SE(3) pose each link actually reaches.
State and action therefore share the same SE(3) language.
We encode the where, when, and what of each world element as tokens,
organize them into a table, and learn a sequence model over this representation.
A pose token at every step
Token
T: 6D pose
Where
t: timestep
When
F: feature
What
Robot action
Robot state
Scene object
Human joint
contextcontextt−1tt+1···t+N−1
Each token is an SE(3) pose paired with a per-trajectory
feature, encoding the where
(T ∈ SE(3)) and the what
(F ∈ ℝD) at each time step. The grid above represents
a dynamic 3D scene as an M × N token tensor: rows index the entity and columns index the
time step. Time indices are
either consecutive within the causal window, or disjoint
virtual indices serving as non-causal context (goal
frames, retargeting exemplars).
Flexible Sequence Modeling
WMM unifies flexible sequence modeling in a single diffusion model, conditioning on any partially known context.
Training
Per-token noise level corrupted tokens
World Motion Model
Predict clean pose tokens(diffusion)
During training, we corrupt each token independently with a
per-token noise level; WMM then predicts the clean tokens.
Inference
At inference, one trained WMM can condition on any partially known context
in the token table and generate the unknown tokens. By changing which tokens are observed,
the same model flexibly supports all the modes below.
(i) Future Prediction P(s′ | s) / Policy P(a | s)
Observe history, predict future tokens: state tokens forecast the scene; action tokens make the same
weights serve as a policy.
NoisyClean
Action
Robot State
Env State
Time →
(ii) Action-Conditioned Prediction P(s′ | s, a)
Keep actions observed and predict future states: WMM becomes a learned forward dynamics simulator,
or world model, that supports planning and model-predictive control (MPC).
NoisyClean
Action
Robot State
Env State
Time →
(iii) Motion Planning / Infilling
Observe sparse keyframes and generate the missing motion between them for planning or interpolation.
NoisyClean
Action
Robot State
Env State
Time →
(iv) Retargeting, Inverse Dynamics P(a | s′, s), etc.
Observe a partial entity set and generate the rest, enabling retargeting, partner prediction, or inverse
dynamics.
NoisyClean
Action
Robot State
Env State
Time →
Network Architecture
Our transformer architecture enhances conditioning on flexible per-token noise levels and scales to thousands
of tokens through factored attention and registers.
•••
•••
•••
Input noised pose tokens
[Opt.] text, image tokens
As prefix
World Motion Model
[a] all
AdaLN
[t] time
AdaLN
[s] reg-t
AdaLN
[p] entity
AdaLN
[q] reg-p
•••
AdaLN
t time, λ noise level, F
feature
registers
token
•••
•••
•••
Output clean pose tokens
Applications
One general architecture, trained the same way, handles all of the tasks below.
WMM can retarget human motion to
humanoid robots (and vice versa).
It models the joint distribution of the human motion and
its physically plausible, retargeted humanoid motion. Most importantly, we do not have to tell WMM
which human joints correspond to which robot joints: a few paired context exemplars suffice, so the
approach is morphology-free. At inference, WMM generates the humanoid motion from
these exemplars and the human motion.
Input 1 · context exemplarsInput 2 · known human motion→Output · generated robot‑link SE(3) trajectories
Inference speed FPS, higher better
GMR
35.0
PHUMA
2.3
Ours
87.0
Tracking success rate higher better
GMR
0.79
PHUMA
0.83
Ours
0.87
Faster and safer retargeting:
compared with optimization-based methods —
GMR
is real-time but less accurate and unsafe;
PHUMA
is more physically plausible but slow and can fail — our data-driven prior, trained on a curated
retargeting dataset, is both faster and safer. We score quality by feeding each kinematic trajectory to a
frozen low-level controller
(TWIST)
and rolling it out in physics simulation.
motion example1 / 2
GMRPHUMAOursKinematics (output)Physics Sim (via TWIST)Input human
WMM can be used to retarget noisy motion reconstructed from videos:
video example1 / 2
Mono Real Video
human·Ours·physics sim
Trained on multiple humanoids:
swap the context exemplar at inference — human+G1 or human+H1
— and the same model retargets different humanoids, each with a different number of links and morphology.
Unitree G1
Unitree H1
TACO — Hand–Object Joint Modeling
WMM also models hand–object interactions.
Using the same joint-modeling formulation as for
retargeting, we train one model on TACO and run it in the different inference modes below:
Future Prediction
hand
object
Time →
known generated
scene 1/5scene 2/5scene 3/5scene 4/5scene 5/5
green first frame is the known input — everything else is generated
Object to Hand
hand
object
Time →
known generated
scene 1/5scene 2/5scene 3/5scene 4/5scene 5/5
green object is the known input — everything else is generated
Hand to Object
hand
object
Time →
known generated
scene 1/5scene 2/5scene 3/5scene 4/5scene 5/5
green hand is the known input — everything else is generated
Motion Planning on Many Morphologies
Motion‑planning inference mask
Arm pose
Arm pose
Obstacle
Time →
known generated
WMM can plan motion for robot
manipulators.
It plans collision‑free motion by modeling the motion of each link
of the robot and the obstacle. Given the first and last robot states together with
the obstacle’s motion in the middle, WMM generates the intervening robot motion.
Input: starting and ending poses, and obstacle motion→Output: planned motion of each link
One model, many morphologies:
a single WMM models robot arms with different morphologies.
PandaUR3UR5UR10WidowXKinovaiiwa14iiwa7
LangTable — WMM as Policy and Forward Dynamics for Model Predictive Control (MPC)
Policy inference mask
Action
Robot State
Env State
Time →
known generated
WMM can act directly as a robot
policy.
It predicts actions alongside future states.
LangTable is a language-conditioned tabletop pushing benchmark. WMM models
the SE(3) state of every block together with the end-effector — its state
(green) and action
(red) — and pins one context time step to the
successful ending frame. At inference, WMM observes only a short history and
generates the rest; we read the policy directly from the predicted
action row.
Input · short history (T=5)Gen 1 (context) · success goalGen 2 · future states + actions
Policy rollout results — please see our paper for quantitative evaluation.
Task: “move the red moon below the
yellow pentagon”
sim renderimagined
goalrollout
Task: “pull the red pentagon apart
from the clump”
sim renderimagined
goalrollout
Task: “push the red pentagon up
and to the left”
sim renderimagined
goalrollout
Task: “put the yellow pentagon to
the bottom right”
sim renderimagined
goalrollout
Task: “slide the green star next
to the red moon”
sim renderimagined
goalrollout
Forward Dynamics mask
Action
Robot State
Env State
Time →
known generated
Box to Absolute Location
Plain policy MPC
success rate (%) ↑
76
80
+4 pp
mean steps ↓
92.7
87.2
−5.5
WMM can also serve as a world
model for planning.
Given the actions as well, the same WMM runs as a forward dynamics simulator and predicts the resulting future
states (mask, right). This unlocks model-predictive control: the model first
imagines a successful goal (below, left), then scores the action chunk’s final frame
against it with a value function based on their difference
and optimizes the chunk with both sampling-based and differentiable gradient-based MPC. This
enables faster rollouts at a higher success rate versus the plain policy (note the rapid push of the
cube in the video below).
Task: “push the blue cube to the top
right”
Value Function (Distance)↔
Imagined GoalPredicted Chunk Last FrameOptimized Action Rollout (note the peak push)
OMOMO — Full-Body Human–Object Joint Modeling
WMM can also model humans
interacting with objects.
Here we demonstrate human–object interaction motion infilling.
Given the first frame of the human and object and
several waypoints of the object’s location, WMM completes the
object’s motion and generates the interacting human motion.
Comparison with CHOIS.
On the OMOMO
dataset, our generation looks more natural than that of CHOIS (ECCV 2024) — even though WMM is not
tailored or optimized specifically for this task. Please see the paper for the quantitative
comparison.
Motion
diversity — one object trajectory, six random seeds. WMM produces varied yet plausible ways
to perform the same object‑conditioned interaction.
TraceGen — Dense 3D Scene Future Prediction
First‑frame inference mask
keypoint 1
keypoint 2
keypoint 400
Time →
first frame (t = 0)
generated, 32 steps
WMM can predict the 3D future of general
scenes.
It handles dense prediction,
querying hundreds of trajectories at once. This formulation subsumes existing
scene-future-prediction works such as
MoMaps,
TraceGen, and
PointWorld.
At inference, the input is the first-frame RGB‑D image; the prediction is the future of the
tracks at dense query positions.
→
Known input: first-frame RGB‑D
Generated output: future 3D tracks
Results
— each scene shows 2D camera tracks (top) and 3D point‑cloud tracks (bottom); the two
columns compare the TraceGen baseline and Ours.
TraceGenOurs
Epic-Kitchens · scene 1DROID · scene 1Epic-Kitchens · scene 2DROID · scene 2
3D‑trajectory error vs. the TraceGen baseline
TraceGen Ours
Droid ↓ lower is better
(×100)
0.206
0.151
MSE
27% improvement
1.289
1.129
MAE
12% improvement
0.285
0.200
Endp.
30% improvement
EpicKitchen ↓ lower is better
(×100)
0.445
0.322
MSE
28% improvement
2.721
2.130
MAE
22% improvement
0.791
0.556
Endp.
30% improvement
Information
BibTeX
@inproceedings{lei2026wmm,
title = {World Motion Models: Flexible Sequence Modeling of SE(3) Trajectories},
author = {Lei, Jiahui and Wang, Qianqian and Darrell, Trevor and Kanazawa, Angjoo},
booktitle = {Advances in Neural Information Processing Systems},
year = {2026}
}
Acknowledgements
This project was funded in part by Meta BAIR partners, the BAIR Humanoid Intelligence Center (BAIR HIC), and
NSF CAREER (No. 2442491).
Capability
One model, many tasks. Click a tile to maximize its interactive viser scene — drag to orbit, scroll
to zoom.
Capability
Methods
How sparse SE(3) trajectories are encoded, sequenced, and conditioned.
Method overview (placeholder). Replace with your real method diagram. Describe inputs, the core
block, conditioning signals, and outputs in one to three sentences.
Results
Detailed quantitative and qualitative results, broken down per application.
A
Application A — title TBA
Replace with the per-application story: dataset, baselines, metrics, and headline numbers. One short
paragraph here; figures and tables below.
Application A. Caption placeholder.B
Application B — title TBA
Per-application narrative for B. Drop numbers in as they land.
Application B. Caption placeholder.C
Application C — title TBA
Per-application narrative for C.
Application C. Caption placeholder.D
Application D — title TBA
Per-application narrative for D.
Application D. Caption placeholder.
Citation & Acknowledgements
BibTeX for the paper, plus thanks to the people and funding that made this work possible.
BibTeX
@article{lei2026wmm,
title = {World Motion Models: Flexible Sequence Modeling of SE(3) Trajectories},
author = {Lei, Jiahui and Wang, Qianqian and Darrell, Trevor and Kanazawa, Angjoo},
journal = {arXiv preprint},
year = {2026}
}
Acknowledgements
We thank our collaborators and funding agencies. Replace this paragraph with the real acknowledgement text
(funding sources, compute providers, helpful discussions).