From Token Probabilities to Semantic Constraints: Towards Declarative Probabilistic Evaluation of Language Models Paper • 2609.13520 • Published 12 days ago
Operadic consistency: a label-free signal for compositional reasoning failures in LLMs Paper • 2606.13649 • Published Jun 11
ArtifactLinker: Linking Scientific Artifacts for Automatic State-of-the-Art Discovery Paper • 2605.16902 • Published May 16 • 1
AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite Paper • 2510.21652 • Published Oct 24, 2025 • 4
TinyScientist: An Interactive, Extensible, and Controllable Framework for Building Research Agents Paper • 2510.06579 • Published Oct 8, 2025
Analytica: Soft Propositional Reasoning for Robust and Scalable LLM-Driven Analysis Paper • 2604.23072 • Published Apr 24
Understanding the Logic of Direct Preference Alignment through Logic Paper • 2412.17696 • Published Dec 23, 2024 • 1