Abstract
A governed analytics framework pairs language models for intent interpretation with deterministic policy execution of pre-approved programs, achieving full answer-and-evidence reliability where runtime-planning agents failed.
We study a governed approach to enterprise analytics: a language model interprets the question, while deterministic policy selects and runs a pre-approved analytical program that returns both results and evidence. We show that this restriction can remain expressive within a defined analytical class, using relational operations plus aggregation, comparison, windows, ranking, and similarity. Fixed meaning, policy, data, and execution rules also make results replayable. Across 440 runs, three 8B models generated SQL and selected tools at runtime, while Qwen3-8B interpreted intent only and policy executed the approved program. None of 330 runtime-planning episodes matched the full answer-and-evidence contract across all test datasets; the policy-executed analyzer matched 110 of 110. This is a configuration-specific result, not evidence that runtime agents cannot succeed under other designs.
Community
The idea is that for supported factual analysis whose approved methods is already known, asking agents to rediscover the method at request time adds a failure surface that is not needed for expressiveness.
For supported enterprise queries, separating language interpretation from policy-selected, reviewed analytical programs is compelling alternative to open-ended tool planning. The model/agent helps determining what the user means without reinventing the measuring procedure.
This is not an argument that agents cannot work, it asks why to introduce runtime method invention when approved method already exists. That is a useful question for anyone building analytical systems people must be able to trust and audit.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- BTS-AgentBench: A Deterministic, Replayable Pipeline from Read-Only Telemetry Logs to Agent Benchmarks (2026)
- Trace Integrity for LLM Data Agents: A Vision for Auditable Structured Reasoning in Real-World Systems (2026)
- SemPlan: Benchmarking Structured Semantic Planning for LLM-Based Queries over Enterprise Data (2026)
- TraceCompiler: Skill-Guided Mining and Compilation of LLM Agent Traces into Mostly Deterministic Workflows (2026)
- BEGIN AI TRANSACTION: Semantic Isolation for Durable AI Workflows (2026)
- Business Truth, not SQL Accuracy: A Rule-Gated 7B Analytics Agent Outperforms a Direct-Prompted 32B Baseline (2026)
- Control Under Compression: Reliability Frontiers for Tool-Using Agents (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.03209 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper