WorldReward: Reward Modeling for Camera-Conditioned World Models
Abstract
WorldReward is a vision-language reward model that evaluates camera-conditioned world models by aligning video chunks with actions and aggregating preferences for both execution consistency and visual quality.
Camera-conditioned world models generate interactive videos in which commanded actions should induce the expected scene changes while appearance, geometry, and temporal dynamics remain coherent. Existing rewards assess these requirements separately: geometry-based rewards estimate trajectory execution but cannot judge the visual quality of the executed motion, whereas image-based rewards measure frame quality without capturing action execution or temporal dynamics. We posit that a vision-language model (VLM) offers a shared reasoning space for relating actions to their visual outcomes. However, judging a complete long video against its full action sequence creates a lengthy, noisy context in which short-lived local action evidence can be missed or diluted. We present WorldReward, a VLM-based pairwise preference reward model that unifies action-consistency and visual-quality evaluation for camera-conditioned world models. WorldReward decomposes paired videos into action-aligned chunks, organizes each chunk into structured visual evidence, and aggregates chunk-level decisions by voting into separate video-level action and visual-quality preferences. To train it, we construct a large-scale reasoning-augmented preference dataset using structured judgments generated by a frontier VLM and refined through tool-based agent auditing and targeted human review. We further introduce WorldReward-Bench, a human-annotated benchmark measuring reward-model agreement with human preferences across action consistency, appearance quality, and motion quality. WorldReward achieves the highest agreement on all three dimensions, exceeding GPT-5.5 by 3.42, 1.45, and 3.56 percentage points, respectively. When used for RL post-training of HY-WorldPlay 1.5, it consistently improves both action execution and visual quality across short- to long-term horizons.
Community
WorldReward: Reward Modeling for Camera-Conditioned World Models
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Sekai2: From World Exploration to Interactive World Modeling (2026)
- VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation (2026)
- AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report (2026)
- WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models (2026)
- EchoWM: Open and Enterable Omnimodal World Models (2026)
- WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity (2026)
- PAVXploreRL: Physical-Action-Visual World Model Reinforcement Learning with Action Exploration (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.03952 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 1
Datasets citing this paper 1
CodeGoat24/WorldReward-Bench
Spaces citing this paper 0
No Space linking this paper