WorldSculpt: Generating Compositional Worlds from Grounded Videos Paper • 2609.05416 • Published 11 days ago • 27
Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning Paper • 2606.00963 • Published May 31
$R^2$-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding Paper • 2404.00801 • Published Mar 31, 2024 • 1
WorldSculpt: Generating Compositional Worlds from Grounded Videos Paper • 2609.05416 • Published 11 days ago • 27
HelloWorld: Enabling Socially Interactive Characters in Video World Models Paper • 2608.05070 • Published Aug 5 • 41
AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report Paper • 2607.18367 • Published Jul 20 • 62
AlayaWorld: Long-Horizon and Playable Video World Generation Paper • 2607.06291 • Published Jul 7 • 93
MolmoAct2: Action Reasoning Models for Real-world Deployment Paper • 2605.02881 • Published May 4 • 357
Affordance-Aware Object Insertion via Mask-Aware Dual Diffusion Paper • 2412.14462 • Published Dec 19, 2024 • 14
Affordance-Aware Object Insertion via Mask-Aware Dual Diffusion Paper • 2412.14462 • Published Dec 19, 2024 • 14
SocialGPT: Prompting LLMs for Social Relation Reasoning via Greedy Segment Optimization Paper • 2410.21411 • Published Oct 28, 2024 • 18