HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published 3 days ago • 231
S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement? Paper • 2608.31100 • Published 4 days ago • 28
AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities Paper • 2607.24821 • Published Jul 17 • 18
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory Paper • 2607.10350 • Published Jul 15 • 85
CLI-Universe: Towards Verifiable Task Synthesis Engine for Terminal Agents Paper • 2606.22883 • Published Jun 22 • 37
CoVEBench: Can Video Editing Models Handle Complex Instructions? Paper • 2606.08415 • Published Jun 7 • 53
TVIR: Building Deep Research Agents Towards Text--Visual Interleaved Report Generation Paper • 2606.02320 • Published Jun 1 • 15
Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization in Agent Trajectories Paper • 2606.02060 • Published Jun 1 • 59
MMG2Skill: Can Agents Distill In-the-Wild Guides into Self-Evolving Skills? Paper • 2606.01993 • Published Jun 1 • 17
OProver: A Unified Framework for Agentic Formal Theorem Proving Paper • 2605.17283 • Published May 17 • 32