Abstract
Modern video games provide a measurable testbed for AI models, combining abilities of visual understanding, instruction decomposition, goal planning, and precise action control over multiple temporal horizons. Existing datasets and benchmarks, however, either cover a narrow range of games, lack language instructions, or rely on high-variance online rollouts. To address these challenges, we introduce GameHorizon, a unified data and evaluation suite that measures gameplay capabilities at different horizons for diverse model families. GameHorizon Suite consists of three components. First, GameHorizon-Annotator is a scalable and automated annotation pipeline for multi-horizon instructions. Second, utilizing the pipeline, we construct GameHorizon-Data, the first large-scale AAA gameplay dataset with temporally aligned videos, player actions, and multi-horizon instructions. It comprises 5,000 hours of recordings from 21 games, collected by 100 human expert players. Third, we build GameHorizon-Bench with reproducible offline and stepwise online testing. The offline track enables reproducible evaluation using thousands of standardized questions organized into three primary tasks and a series of diagnostic variants, while the online track tests whether offline scores reflect actual gameplay capabilities and localizes failures to specific steps within long-horizon gameplay. Based on our GameHorizon Suite, we evaluate 47 models through more than one million model invocations, revealing a meaningful hierarchy of task difficulty and pronounced differences in model capabilities. Our work can provide a standardized yardstick for evaluating gameplay capabilities across horizons and model families. We will release our dataset, annotator, and benchmark to facilitate future research.
Community
We introduce GameHorizon, a large-scale data and evaluation suite spanning multiple temporal horizons and diverse AAA games. It serves as a standardized yardstick across a broad range of model families, including VLMs, UMMs, GUI agents, coding agents, and game agents. We will release our dataset, annotator, and benchmark to facilitate future research.
Github:https://github.com/TencentARC/GameHorizon
Project Page:https://gamehorizon-suite.github.io
Arxiv:https://github.com/TencentARC/GameHorizon
This is an automated message from the ResearchStudio team.
We created an interactive ResearchStudio Reel for this paper. It includes a visual poster, a video, and a blog, all available for download in editable formats.
Open the ResearchStudio Reel →
Download all files from Hugging Face
Please give this comment a thumbs up if you find the Reel helpful!
Want to explore or create Reels for more papers? Visit the ResearchStudio demo.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- GameWAM: A World Action Model for Video Games (2026)
- StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation (2026)
- GameXpert-Bench: How Far Are Coding Agents from Expert Game Development? (2026)
- SocialReasonBench: A Video-QA Benchmark for Social Reasoning with Counterfactual Narrative Videos (2026)
- PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives (2026)
- Towards Comprehensive Basketball Understanding (2026)
- GameASG-Bench: Benchmarking Autonomous Software Generation for Game Development (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.25001 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
