arXiv:2608.22538cs.AI2026-08

将政策推理与确定性流程控制分离,提升复杂任务的执行成功率和稳定性。

STAGE: Stateful Translation to Agentic Graph Execution with Policy-Scoped Context and Deterministic Control

论文配图:STAGE: Stateful Translation to Agentic Graph Execution with Policy-Scoped Context and Deterministic Control
图 1 · 摘自论文原文
  • 在政策节点内进行局部推理,流程控制由确定性代码实现。
  • 在深层工作流中任务成功率最高提升65.7个百分点。
  • 适合需要高可靠性的企业级自动化决策场景。

受政策约束的智能体需在遵循授权流程的同时解读案件证据。我们提出STAGE,一种可执行图框架,将模型判断限制在政策作用域节点内,同时将流程控制交由确定性代码。我们在三个公开政策遵循基准和一个专有银行基准Smart Dispute上评估了STAGE。相比整体政策执行,STAGE提升了任务成功率与重复运行的可靠性,尤其在深层工作流中表现突出:在$τ^2$-bench Telecom和Smart Dispute上,$ ext{Pass}^{3}$分别提升最高达55.0和65.7个百分点。结果表明,将局部政策推理与确定性流程控制结合,对企事业单位具有显著价值。

原文摘要 · Abstract (English)

Policy-governed agents must interpret case evidence while reliably following authorized procedures. We present STAGE, an executable-graph framework that confines model judgment to policy-scoped nodes while placing procedural control in deterministic code. We evaluate STAGE on three public policy-following benchmarks and Smart Dispute, a proprietary banking benchmark. Compared with monolithic full-policy execution, STAGE improves task success and repeated-run reliability, with its largest observed gains on the deeper workflows. On $τ^2$-bench Telecom and Smart Dispute, $\mathrm{Pass}^{3}$ improves by up to 55.0 and 65.7 percentage points, respectively. These results demonstrate the value of combining localized policy reasoning with deterministic procedural control for enterprise use.

智能体流程控制政策遵循可执行图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。