Collections
Discover the best community collections!
Collections including paper arxiv:2609.10522
-
Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization
Paper • 2608.26103 • Published • 26 -
Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models
Paper • 2608.27550 • Published • 95 -
Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning
Paper • 2608.27549 • Published • 53 -
Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models
Paper • 2608.25518 • Published • 196
-
IRASim: Learning Interactive Real-Robot Action Simulators
Paper • 2406.14540 • Published • 6 -
WildActor: Unconstrained Identity-Preserving Video Generation
Paper • 2603.00586 • Published • 38 -
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data
Paper • 2605.18287 • Published • 15 -
PhysiFormer: Learning to Simulate Mechanics in World Space
Paper • 2606.27364 • Published • 12
-
LLM Pruning and Distillation in Practice: The Minitron Approach
Paper • 2408.11796 • Published • 62 -
TableBench: A Comprehensive and Complex Benchmark for Table Question Answering
Paper • 2408.09174 • Published • 53 -
To Code, or Not To Code? Exploring Impact of Code in Pre-training
Paper • 2408.10914 • Published • 45 -
Open-FinLLMs: Open Multimodal Large Language Models for Financial Applications
Paper • 2408.11878 • Published • 64
-
World Value Models for Robotic Manipulation
Paper • 2606.24742 • Published • 8 -
In-Context World Modeling for Robotic Control
Paper • 2606.26025 • Published • 64 -
ASPIRE: Agentic /Skills Discovery for Robotics
Paper • 2607.00272 • Published • 28 -
τ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation
Paper • 2608.16885 • Published • 16
-
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective
Paper • 2507.01925 • Published • 39 -
Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
Paper • 2507.16746 • Published • 36 -
MolmoAct: Action Reasoning Models that can Reason in Space
Paper • 2508.07917 • Published • 45 -
Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies
Paper • 2508.20072 • Published • 32
-
Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization
Paper • 2608.26103 • Published • 26 -
Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models
Paper • 2608.27550 • Published • 95 -
Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning
Paper • 2608.27549 • Published • 53 -
Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models
Paper • 2608.25518 • Published • 196
-
World Value Models for Robotic Manipulation
Paper • 2606.24742 • Published • 8 -
In-Context World Modeling for Robotic Control
Paper • 2606.26025 • Published • 64 -
ASPIRE: Agentic /Skills Discovery for Robotics
Paper • 2607.00272 • Published • 28 -
τ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation
Paper • 2608.16885 • Published • 16
-
IRASim: Learning Interactive Real-Robot Action Simulators
Paper • 2406.14540 • Published • 6 -
WildActor: Unconstrained Identity-Preserving Video Generation
Paper • 2603.00586 • Published • 38 -
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data
Paper • 2605.18287 • Published • 15 -
PhysiFormer: Learning to Simulate Mechanics in World Space
Paper • 2606.27364 • Published • 12
-
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective
Paper • 2507.01925 • Published • 39 -
Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
Paper • 2507.16746 • Published • 36 -
MolmoAct: Action Reasoning Models that can Reason in Space
Paper • 2508.07917 • Published • 45 -
Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies
Paper • 2508.20072 • Published • 32
-
LLM Pruning and Distillation in Practice: The Minitron Approach
Paper • 2408.11796 • Published • 62 -
TableBench: A Comprehensive and Complex Benchmark for Table Question Answering
Paper • 2408.09174 • Published • 53 -
To Code, or Not To Code? Exploring Impact of Code in Pre-training
Paper • 2408.10914 • Published • 45 -
Open-FinLLMs: Open Multimodal Large Language Models for Financial Applications
Paper • 2408.11878 • Published • 64