DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
Abstract
DreamX-Phi 1.0 is an action-conditioned video world model for robotic manipulation that uses geometric attention encoding, depth estimation, object masks with a frozen teacher, and distillation to generate faithful future observations.
We present DreamX-Phi 1.0, an action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and a prescribed action sequence comprising end-effector poses and gripper states, predicts the resulting future observations. Yet realism alone does not guarantee faithfulness: a convincing rollout can still move the wrong arm or lose the manipulated object. To ensure the prediction respects each arm's commanded path, we inject per-arm SE(3) transformations into attention via PRoPE-style geometric encoding, preserving arm identity and rigid-motion structure. Action control alone does not fully constrain scene geometry or the evolution of small manipulated objects. We therefore add a lightweight depth branch for scene-level geometry and use SAM3 masks with a frozen V-JEPA teacher to maintain object consistency throughout grasping. We further distill the multi-step generator into a few-step student via distribution-matching distillation for efficient deployment. At the time of writing, achieves first place on Track~1 and second place on Track~2 of the WorldArena~2.0 Challenge. Our model and code will be publicly available.
Community
At the time of writing, DreamX-Phi 1.0 achieves first place on Track 1 and second place on Track 2 of the WorldArena 2.0 Challenge. Model weights and inference code will be made publicly available (https://github.com/AMAP-ML/DreamX-Phi) after the WorldArena 2.0 IROS Challenge concludes.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE (2026)
- Foresight Without Seeing: Latent Futures for World Action Models (2026)
- WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN (2026)
- World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation (2026)
- Native Video-Action Pretraining for Generalizable Robot Control (2026)
- JEPA-WAM: Stage-Level Joint-Embedding Prediction for World-Action Models in Robot Manipulation (2026)
- ContactFlow: A video action conditioning that transfers across embodiments (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.13489 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper