arXiv:2604.18934cs.AI2026-04被引 1

测试AI代理在跨应用自动化中的真实能力,涵盖找接口、守规则、跨系统协作。

AutomationBench

论文配图:AutomationBench
图 1 · 摘自论文原文
  • 模拟真实业务流程,要求代理自主发现API并协调多系统
  • 当前顶尖模型在复杂任务上得分不足10%
  • 适合评估AI代理在企业级自动化中的实用性

现有软件自动化基准很少同时包含跨应用协同、自主API发现和策略合规性。真实业务流程涉及CRM、收件箱、日历、消息平台等多个系统,需代理自行发现正确端点,遵守分层业务规则,并向各系统准确写入数据。为此,我们提出AutomationBench,一个基于Zapier平台真实工作流模式的基准,覆盖销售、营销、运营、支持、财务和人力资源领域。代理需自主发现相关端点,遵循复杂策略,应对无关或误导性记录。评分采用程序化端态判断:数据是否正确落入目标系统。目前最先进模型得分低于10%。AutomationBench为评估当前模型在企业所需智能体能力上的真实水平提供了挑战性且真实的衡量标准。

原文摘要 · Abstract (English)

Existing AI benchmarks for software automation rarely combine cross-application coordination, autonomous API discovery, and policy adherence. Real business workflows demand all three: a single task may span a CRM, inbox, calendar, and messaging platform - requiring the agent to find the right endpoints, follow a policy document, and write correct data to each system. To address this gap, we introduce AutomationBench, a benchmark for evaluating AI agents on cross-application workflow orchestration via REST APIs. Drawing on real workflow patterns from Zapier's platform, tasks span Sales, Marketing, Operations, Support, Finance, and HR domains. Agents must discover relevant endpoints themselves, follow layered business rules, and navigate environments with irrelevant and sometimes misleading records. Grading is programmatic and end-state only: whether the correct data ended up in the right systems. Even the best frontier models currently score below 10%. AutomationBench provides a challenging, realistic measure of where current models stand relative to the agentic capabilities businesses actually need.

AI代理自动化基准测试跨系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。