评估AI代理在企业部署中的可靠性与成本,找出可行的监督方案。
READY or Not: Reliable Enterprise Agent Deployment
- 通过统一流程评估代理在真实工作流中的可靠性和开销。
- 同一可靠性目标下,不同系统需的人工审查比例相差达9.6%。
- 适用于医疗审计等高风险场景,支持量化决策和监督策略制定。
AI代理在基准测试中表现良好,仍可能不适合企业部署。现有基准衡量代理完成专业任务的能力,而企业部署关注的是:在可接受的人工监督下,是否能达到指定可靠性水平且成本可控。本文提出可靠企业代理部署框架(READY),在保留各工作流自身成功定义的同时,采用统一资格评估流程。给定代理、工作流及候选监督策略,READY测量人机系统的可靠性与运营成本,选择满足可靠性目标的最低成本策略,并对未见案例进行统计资格认证。最终输出部署配置:可靠性、人工负担与成本。该框架作为开源测试平台,解耦工作流定义、执行、评估与认证,运行于现有代理评估基础设施。在涵盖16个代理系统、750个案例的临床审计全流程研究中,发现两个自主准确率仅差0.3个百分点(72.8% vs. 72.5%)的系统,在相同76%可靠性目标下,所需人工审查比例分别为39.2%与29.6%,差异显著。因此,READY将企业代理评估从‘能做多好’转向‘在何种条件与成本下可可靠部署’。通过明确并可统计验证部署条件,为代理系统比较、监督标准设定与证据驱动决策提供基础。
原文摘要 · Abstract (English)
An AI agent can perform well on benchmarks and still be unsuitable for deployment. Existing AI-agent benchmarks measure whether an agent can complete realistic professional work, whereas enterprise deployment asks a different question: whether an agent can meet a required reliability level, under acceptable human oversight, and at tolerable cost. We introduce Reliable Enterprise Agent Deployment (READY), a framework for qualifying AI agents for deployment on enterprise workflows. READY preserves each workflow's own definition of successful execution while applying a common qualification procedure. Given an agent, a workflow, and a class of candidate oversight policies, READY measures the reliability and operating cost of the human-AI system, selects the minimum-cost policy that satisfies a specified reliability target, and statistically qualifies it on held-out cases. The resulting deployment profile characterizes the supported operating point: reliability, human-oversight burden, and cost. READY is implemented as an open testbed that decouples workflow specification, execution, evaluation, and qualification, and runs on existing agent-evaluation infrastructure. In an end-to-end clinical-audit case study spanning 16 agent systems and 750 cases, READY reveals differences hidden by autonomous performance: two systems separated by only 0.3 points in autonomous accuracy (72.8% vs. 72.5%) require 39.2% versus 29.6% human review, respectively, to qualify at the same 76% reliability target under the evaluated oversight policy. READY thus shifts enterprise agent evaluation from how well can the agent perform the work? to under what conditions, and at what cost, can it be reliably deployed? By making those conditions explicit and statistically testable, READY provides a basis for comparing agent systems, setting oversight requirements, and making evidence-based deployment decisions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。