用确定性策略执行预批准分析程序,确保结果可复现且准确。
MasterControl Seventeen Every Time
- 语言模型理解问题,确定性策略选择并运行预审批程序。
- 440次测试中,策略执行方案110/110匹配完整答案与证据。
- 适合对结果可复现性要求高的企业级分析场景。
我们研究了一种受控的企业分析方法:语言模型解析问题,确定性策略选择并运行预批准的分析程序,返回结果与证据。该方法在关系操作、聚合、比较、窗口、排序和相似性等分析任务中保持表达力。固定语义、策略、数据与执行规则使结果具备可复现性。在440次运行中,三个8B模型实时生成SQL并选择工具,而Qwen3-8B仅解析意图,策略执行预审批程序。330次实时规划未在所有测试数据集上达成完整答案与证据契约;策略执行的分析器则全部匹配(110/110)。此为特定配置下的结果,不证明其他设计下运行时智能体无法成功。
原文摘要 · Abstract (English)
We study a governed approach to enterprise analytics: a language model interprets the question, while deterministic policy selects and runs a pre-approved analytical program that returns both results and evidence. We show that this restriction can remain expressive within a defined analytical class, using relational operations plus aggregation, comparison, windows, ranking, and similarity. Fixed meaning, policy, data, and execution rules also make results replayable. Across 440 runs, three 8B models generated SQL and selected tools at runtime, while Qwen3-8B interpreted intent only and policy executed the approved program. None of 330 runtime-planning episodes matched the full answer-and-evidence contract across all test datasets; the policy-executed analyzer matched 110 of 110. This is a configuration-specific result, not evidence that runtime agents cannot succeed under other designs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。