arXiv:2509.26506cs.AI2025-09被引 7

SCUBA benchmark测试企业级软件自动化能力,揭示大模型在真实工作流中的表现差距。

SCUBA: Salesforce Computer Use Benchmark

  • 基于真实用户访谈构建300个销售平台任务,覆盖三类角色的完整工作流。
  • 开源模型零样本成功率不足5%,闭源模型最高达39%,演示增强后可提升至50%。
  • 支持沙箱环境并行执行,提供细粒度评估,适合研究企业自动化智能体者。

我们提出SCUBA,一个用于评估计算机使用智能体在Salesforce平台客户关系管理(CRM)工作流中表现的基准。该基准包含300个源自真实用户访谈的任务实例,覆盖平台管理员、销售代表和服务代理三类主要角色。任务涵盖企业级关键能力:企业软件界面导航、数据操作、流程自动化、信息检索与故障排查。为保证真实性,SCUBA在Salesforce沙箱环境中运行,支持并行执行及细粒度评估指标以捕捉里程碑进展。我们在零样本和示范增强两种设置下对多种智能体进行基准测试。结果显示不同智能体范式间存在巨大性能差距,且开源模型与闭源模型之间差距显著:在零样本设置下,仅在类似OSWorld等基准上表现良好的开源模型,其在SCUBA上的成功率低于5%;而基于闭源模型的方法仍可达39%的成功率。在示范增强设置下,任务成功率提升至50%,同时时间与成本分别降低13%和16%。这些发现凸显了企业任务自动化的挑战与智能体方案的潜力。通过提供一个具现实性的基准与可解释的评估体系,SCUBA旨在加速复杂商业软件生态中可靠计算机使用智能体的发展。

原文摘要 · Abstract (English)

We introduce SCUBA, a benchmark designed to evaluate computer-use agents on customer relationship management (CRM) workflows within the Salesforce platform. SCUBA contains 300 task instances derived from real user interviews, spanning three primary personas, platform administrators, sales representatives, and service agents. The tasks test a range of enterprise-critical abilities, including Enterprise Software UI navigation, data manipulation, workflow automation, information retrieval, and troubleshooting. To ensure realism, SCUBA operates in Salesforce sandbox environments with support for parallel execution and fine-grained evaluation metrics to capture milestone progress. We benchmark a diverse set of agents under both zero-shot and demonstration-augmented settings. We observed huge performance gaps in different agent design paradigms and gaps between the open-source model and the closed-source model. In the zero-shot setting, open-source model powered computer-use agents that have strong performance on related benchmarks like OSWorld only have less than 5\% success rate on SCUBA, while methods built on closed-source models can still have up to 39% task success rate. In the demonstration-augmented settings, task success rates can be improved to 50\% while simultaneously reducing time and costs by 13% and 16%, respectively. These findings highlight both the challenges of enterprise tasks automation and the promise of agentic solutions. By offering a realistic benchmark with interpretable evaluation, SCUBA aims to accelerate progress in building reliable computer-use agents for complex business software ecosystems.

智能体企业自动化基准测试Salesforce

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。