Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning
Paper • 2608.02831 • Published • 13
None defined yet.
How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review
Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use