WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory
Abstract
Video world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations over long horizons and across viewpoints. We present WorldCrafter, a video world model that learns a camera-queryable implicit 3D-aware memory for this purpose. The key insight is to let the requested viewpoint shape how multi-view evidence is compressed into the video generator's limited token budget. Trained jointly with the video generator, a memory encoder and pose-conditioned readout module integrate historical observations into a fixed set of target view-specific tokens before denoising, without explicit depth-based correspondences. By combining this memory with recent temporal context and few-step distillation, WorldCrafter enables streaming scene exploration from a single input image or text prompt. Experiments across static and dynamic scenes show substantial gains in long-horizon consistency and camera-control accuracy while preserving visual quality during minute-scale exploration.
Community
We introduce WorldCrafter, a video world model with a camera-queryable implicit 3D-aware memory for consistent, long-horizon scene exploration. Starting from a single image or text prompt, WorldCrafter combines historical observations, recent temporal context, and few-step distillation to support streaming exploration with improved consistency and camera control.
Project page: https://drexubery.github.io/WorldCrafter
This is an automated message from the ResearchStudio team.
We created an interactive ResearchStudio Reel for this paper. It includes a visual poster, a video, and a blog, all available for download in editable formats.
Open the ResearchStudio Reel โ
Download all files from Hugging Face
Please give this comment a thumbs up if you find the Reel helpful!
Want to explore or create Reels for more papers? Visit the ResearchStudio demo.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Wonder: Video World Model Done Better (2026)
- AlayaVista: Streaming World Modeling from Panoramic States to Perspective Video (2026)
- Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory (2026)
- MiniWorld: Democratizing the Training of Video World Models from Scratch (2026)
- World in World: Explore the World with World Models (2026)
- JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion (2026)
- LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.24984 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 2
TencentARC/WorldCrafter-Base
Datasets citing this paper 0
No dataset linking this paper
