AI & ML interests
None defined yet.
Recent Activity
View all activity
Papers
ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
datasets 147
benchflow/frontierphysics-traj
Updated • 823 • 1
benchflow/frontierphysics-pr552-evidence
Updated
benchflow/frontierphysics-pr586-evidence
Updated
benchflow/frontierphysics-pr51-evidence
Updated
benchflow/frontierphysics-pr575-evidence
Updated
benchflow/frontierphysics-pr437-evidence
Updated • 14
benchflow/frontierphysics-pr50-evidence
Updated
benchflow/frontierphysics-pr25-evidence
Viewer • Updated • 1.67k • 69
benchflow/frontierphysics-pr65-evidence
Viewer • Updated • 412 • 21
benchflow/frontierphysics-pr38-evidence
Viewer • Updated • 272 • 58