CLI-agent trajectories on ALE-Bench (arXiv 2506.09050): ahc026+ahc039 lite, raw streams + judged submissions + final results.
Agent Native Research Lab
non-profit
AI & ML interests
None defined yet.
Recent Activity
View all activity
Phase-2 evals: agents playing from token-truncated ARAs, one repo per game.
Live Hard Task Bench trajectories, one repo per model.
-
AgentNativeResearchLab/lhtb-gemini3.1-pro-high-trajectories
Preview • Updated • 94 -
AgentNativeResearchLab/lhtb-glm5.2-trajectories
Preview • Updated • 176 -
AgentNativeResearchLab/lhtb-gpt5.6-sol-trajectories
Preview • Updated • 344 • 1 -
AgentNativeResearchLab/lhtb-grok4.5-trajectories
Preview • Updated • 138
CLI-agent trajectories on ALE-Bench (arXiv 2506.09050): ahc026+ahc039 lite, raw streams + judged submissions + final results.
Full record per harness×model×game: trajectories, recordings, replays + the ARA the agent built live. Same game across models = comparison unit.
Phase-2 evals: agents playing from token-truncated ARAs, one repo per game.
ARAs from agents solving DiscoverPhysics, one repo per model.
Live Hard Task Bench trajectories, one repo per model.
-
AgentNativeResearchLab/lhtb-gemini3.1-pro-high-trajectories
Preview • Updated • 94 -
AgentNativeResearchLab/lhtb-glm5.2-trajectories
Preview • Updated • 176 -
AgentNativeResearchLab/lhtb-gpt5.6-sol-trajectories
Preview • Updated • 344 • 1 -
AgentNativeResearchLab/lhtb-grok4.5-trajectories
Preview • Updated • 138
Foundation-model open-problems trajectories, one repo per model.