Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs Paper • 2608.20492 • Published 12 days ago • 110
AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces Paper • 2608.23041 • Published 8 days ago • 63
CAFE: Self-Improving Search Agents Need Co-Evolving Feedback Paper • 2608.24794 • Published 7 days ago • 6