From Pixels to States: Rethinking Interactive World Models as Game Engines Paper • 2607.14076 • Published 6 days ago • 33
Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos Paper • 2607.11523 • Published 8 days ago • 13
SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Paper • 2607.08688 • Published 12 days ago • 9
KVpop -- Key-Value Cache Compression with Predictive Online Pruning Paper • 2607.05061 • Published 15 days ago • 23
AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents Paper • 2607.02255 • Published 19 days ago • 64
OcclusionFormer: Arranging Z-Order for Layout-Grounded Image Generation Paper • 2605.21343 • Published May 20 • 8
PanoWorld: Towards Spatial Supersensing in 360^circ Panorama World Paper • 2605.13169 • Published May 13 • 21
AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation Paper • 2605.13724 • Published May 13 • 105
Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization Paper • 2605.10780 • Published May 12 • 33