CaML Benchmarks Collection CaML benchmark datasets: TAC, ANIMA, MORU, WAS + holdout/audit sets. Leaderboard: compassionbench.com • 10 items • Updated 2 days ago
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation Paper • 2607.15434 • Published 19 days ago • 4
Do LLMs Hold Their Values? MANTA: A Multi-Turn Adversarial Benchmark for Animal Welfare Reasoning Paper • 2605.16301 • Published Jun 3 • 1
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation Paper • 2607.15434 • Published 19 days ago • 4
Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models Paper • 2606.18142 • Published Jun 17 • 2
CompassioninMachineLearning/SDF_docs_animals_20260718_1735_sonnetpilot Viewer • Updated 20 days ago • 182 • 67