ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction Paper • 2607.29677 • Published 6 days ago • 20
LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger Paper • 2607.28374 • Published 7 days ago • 13
Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability Paper • 2607.26637 • Published 8 days ago • 12
MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Paper • 2607.25948 • Published 9 days ago • 18
Pass the Baton: Trajectory-Relayed On-Policy Distillation Paper • 2607.26057 • Published 9 days ago • 33
HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone Paper • 2607.25895 • Published 9 days ago • 156
DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes Paper • 2607.24516 • Published 10 days ago • 7
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD Paper • 2607.20145 • Published 15 days ago • 74
Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers Paper • 2607.19139 • Published 16 days ago • 74
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Paper • 2607.14952 • Published 21 days ago • 208
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU Paper • 2607.19191 • Published 16 days ago • 309
EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World Paper • 2607.17250 • Published 18 days ago • 92
SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning Paper • 2607.10966 • Published 24 days ago • 5
S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation Paper • 2607.15686 • Published 20 days ago • 16
KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation Paper • 2607.14202 • Published 22 days ago • 42
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation Paper • 2607.11886 • Published 24 days ago • 84
Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models Paper • 2607.12463 • Published 23 days ago • 107
SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding Paper • 2607.10400 • Published 26 days ago • 71
ABot-N1: Toward a General Visual Language Navigation Foundation Model Paper • 2607.10383 • Published 23 days ago • 102