arXiv:2603.24582cs.AI2026-03

提出可量化智能体决策可靠性的概率框架,助力企业审计与成本控制。

The Stochastic Gap: A Markovian Framework for Pre-Deployment Reliability and Oversight-Cost Auditing in Agentic Artificial Intelligence

  • 构建基于马尔可夫过程的可靠性评估框架,量化决策盲区与监督成本。
  • 在采购流程中发现:状态扩展后盲区从1.65%升至12.53%,显著影响自主性可信度。
  • 适合关注企业AI治理、流程自动化与风险审计的技术团队使用。

组织中的智能体人工智能是受可靠性与监督成本约束的序列决策问题。当确定性流程被动作和工具调用的随机策略取代时,关键问题不再是下一步是否合理,而是整个决策轨迹是否在统计上成立、局部清晰且经济可控。本文建立了一套测度论意义上的马尔可夫框架,核心量包括状态盲点质量 B_n(tau)、状态-动作盲质量 B^SA_{pi,n}(tau)、基于熵的人工介入升级机制,以及工作流访问测度上的期望监督成本恒等式。我们在2019年业务流程智能挑战赛的采购付款日志(251,734个案例,1,595,923个事件,42种工作流动作)上实例化该框架,并基于时间顺序80/20划分构建了一个日志驱动的模拟智能体。主要实证发现为:即使工作流在状态层面看似稳健,其下一步决策仍存在显著盲区;将操作状态扩展至包含案件上下文、经济规模和行为者类别后,状态空间由42增至668,状态-动作盲质量从tau=50时的0.0165上升至tau=1000时的0.1253。在保留测试集上,m(s) = max_a pi-hat(a|s) 平均仅偏差实际自主步骤准确率3.4个百分点。决定统计可信自主性的同一组量也决定了预期监督负担。该框架已在大规模企业采购流程中验证,适用于有运营事件日志的工程流程直接应用。

原文摘要 · Abstract (English)

Agentic artificial intelligence (AI) in organizations is a sequential decision problem constrained by reliability and oversight cost. When deterministic workflows are replaced by stochastic policies over actions and tool calls, the key question is not whether a next step appears plausible, but whether the resulting trajectory remains statistically supported, locally unambiguous, and economically governable. We develop a measure-theoretic Markov framework for this setting. The core quantities are state blind-spot mass B_n(tau), state-action blind mass B^SA_{pi,n}(tau), an entropy-based human-in-the-loop escalation gate, and an expected oversight-cost identity over the workflow visitation measure. We instantiate the framework on the Business Process Intelligence Challenge 2019 purchase-to-pay log (251,734 cases, 1,595,923 events, 42 distinct workflow actions) and construct a log-driven simulated agent from a chronological 80/20 split of the same process. The main empirical finding is that a large workflow can appear well supported at the state level while retaining substantial blind mass over next-step decisions: refining the operational state to include case context, economic magnitude, and actor class expands the state space from 42 to 668 and raises state-action blind mass from 0.0165 at tau=50 to 0.1253 at tau=1000. On the held-out split, m(s) = max_a pi-hat(a|s) tracks realized autonomous step accuracy within 3.4 percentage points on average. The same quantities that delimit statistically credible autonomy also determine expected oversight burden. The framework is demonstrated on a large-scale enterprise procurement workflow and is designed for direct application to engineering processes for which operational event logs are available.

AI治理流程审计马尔可夫模型监督成本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。