SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research? Paper • 2609.09113 • Published 9 days ago • 20
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows Paper • 2608.17800 • Published 30 days ago • 11
SWE-Touch: Benchmarking Coding Agents When Users Touch the Code Paper • 2608.02499 • Published Aug 3 • 25
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows Paper • 2608.17800 • Published 30 days ago • 11
Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application Paper • 2606.12191 • Published Jun 10 • 73
SWE-Touch: Benchmarking Coding Agents When Users Touch the Code Paper • 2608.02499 • Published Aug 3 • 25
HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone Paper • 2607.25895 • Published Jul 28 • 160
Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity Paper • 2607.00248 • Published Jun 30 • 33
DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios Paper • 2604.25914 • Published Apr 28 • 43 • 5
GATE: Graph-based Adaptive Tool Evolution Across Diverse Tasks Paper • 2502.14848 • Published Feb 20, 2025 • 1
DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios Paper • 2604.25914 • Published Apr 28 • 43
DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios Paper • 2604.25914 • Published Apr 28 • 43 • 5
DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios Paper • 2604.25914 • Published Apr 28 • 43