GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Paper • 2608.05747 • Published 2 days ago • 36
Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data Paper • 2608.02580 • Published 5 days ago • 22
Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing Paper • 2608.02711 • Published 5 days ago • 80
GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience Paper • 2608.02392 • Published 5 days ago • 14
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Paper • 2608.02023 • Published 5 days ago • 152
DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF Image-Text-to-Text • 9B • Updated 1 day ago • 343k • 305
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published 9 days ago • 302