Dr. Claw: An AI Scientist Workspace for Vibe Research Paper • 2609.00365 • Published 14 days ago • 179
Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them? Paper • 2609.10226 • Published 5 days ago • 21
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published 6 days ago • 412
PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback Paper • 2608.30241 • Published 14 days ago • 12
Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models Paper • 2608.25518 • Published 19 days ago • 196
MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models Paper • 2607.27637 • Published Aug 1 • 6
TARS: Timestep-Aware Data Scaling for 3D-Free Video Re-Shooting Paper • 2607.28261 • Published Jul 30 • 116
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published Jul 30 • 311
DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation Paper • 2607.26811 • Published Jul 29 • 93
Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs Paper • 2607.00302 • Published Jul 1 • 2
LLM-as-a-Verifier: A General-Purpose Verification Framework Paper • 2607.05391 • Published Jul 6 • 19
Lexical Consensus: Grounded Word Learning and Shared Meaning in Artificial Agents Paper • 2606.22207 • Published Jun 20 • 4
Retrospective Harness Optimization: Improving LLM Agents via Self-Preference over Trajectory Rollouts Paper • 2606.05922 • Published Jun 4 • 72
TVIR: Building Deep Research Agents Towards Text--Visual Interleaved Report Generation Paper • 2606.02320 • Published Jun 1 • 15
Convex Low-resource Accent-Robust Language Detection in Speech Recognition Paper • 2605.23235 • Published May 22 • 6
DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards Paper • 2605.21467 • Published May 20 • 207
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization Paper • 2605.17757 • Published May 18 • 66
Spreadsheet-RL: Advancing Large Language Model Agents on Realistic Spreadsheet Tasks via Reinforcement Learning Paper • 2605.22642 • Published May 21 • 36