view article Article How UK AISI and EvalEval Are Making Benchmark Results Reproducible +7 evijit, j-chim, deeplumiere, srishtiy, wmmkennedy, irenesolaiman, mcfadyen-aisi, lynn-aisi, coz-aisi • about 23 hours ago • 4
view article Article State of Open Models: Summer 2026 Observations +1 AdinaY, multimodalart, irenesolaiman • Aug 14 • 213
view article Article Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident +2 hlarcher, XciD, raphael-gl, chris-rannou • Jul 27 • 504
view article Article UN Remarks on "Navigating the Challenges of AI in Cyberspace" 🌐 yjernite • Jul 22 • 3
view article Article Featuring Every Eval Ever Results on Hugging Face Model Pages +5 deepmage121, evijit, SaylorTwift, janbatzner, borgr, irenesolaiman, julien-c • Jun 30 • 54
Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results Paper • 2606.14516 • Published Jun 12 • 7
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Paper • 2606.09809 • Published Jun 8 • 5
view article Article Introducing Evaluation Cards: A Live Interpretive Layer for Understanding the AI Evaluations Ecosystem evaleval • Jun 11 • 1
LLM-42: Enabling Determinism in LLM Inference with Verified Speculation Paper • 2601.17768 • Published Jan 30 • 1
The Silent Hyperparameter: Quantifying the Impact of Inference Backends on LLM Reproducibility Paper • 2605.19537 • Published May 20 • 2
view article Article LeRobot Humanoid: An Open, Low-Cost, 3D-Printed Humanoid for Robot Learning VirgileBatto • May 21 • 72
view article Article Harness, Scaffold, and the AI Agent Terms Worth Getting Right sergiopaniego, ariG23498 • May 25 • 148