PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives Paper • 2608.13552 • Published 6 days ago • 44
The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images Paper • 2608.06270 • Published 13 days ago • 7
InSight-doc Collection Agentic Visual Perception for Long-Document Understanding • 5 items • Updated 3 days ago • 1
InSight-doc: Agentic Visual Perception for Long-Document Understanding Paper • 2608.10628 • Published 8 days ago • 11
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation Paper • 2607.11886 • Published Jul 13 • 85
MM-Zero: Self-Evolving Multi-Model Vision Language Models From Zero Data Paper • 2603.09206 • Published Mar 10 • 54
AgentVista: Evaluating Multimodal Agents in Ultra-Challenging Realistic Visual Scenarios Paper • 2602.23166 • Published Feb 26 • 45
VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning? Paper • 2603.07888 • Published Mar 9 • 10
OoD-Bench: Quantifying and Understanding Two Dimensions of Out-of-Distribution Generalization Paper • 2106.03721 • Published Jun 7, 2021 • 1
CODA: A Real-World Road Corner Case Dataset for Object Detection in Autonomous Driving Paper • 2203.07724 • Published Mar 15, 2022 • 1
Dual Risk Minimization: Towards Next-Level Robustness in Fine-tuning Zero-Shot Models Paper • 2411.19757 • Published Nov 29, 2024 • 1
MapTrace: Scalable Data Generation for Route Tracing on Maps Paper • 2512.19609 • Published Dec 22, 2025 • 3
InSight-o3 Collection Empowering Multimodal Foundation Models with Generalized Visual Search • 5 items • Updated Mar 24 • 1
InSight-o3: Empowering Multimodal Foundation Models with Generalized Visual Search Paper • 2512.18745 • Published Dec 21, 2025 • 12