评测企业级智能体在复杂工作流中的规划能力,发现顶尖模型成功率仅37.4%。
EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings
- 构建包含164张表和512个工具的模拟企业环境,评估智能体长期规划与工具使用能力。
- 14个前沿模型在1150项任务中平均成功率不足四成,最优模型仅达37.4%。
- 揭示当前智能体缺乏合理拒绝能力,易产生有害后果,不适合自主部署。
大型语言模型正从被动信息提供者转向执行复杂工作流的主动智能体。然而,其在企业环境中作为可靠AI员工的部署受阻于现有基准无法体现专业场景的复杂性,特别是长期规划中持续的状态变化与严格的访问协议需求。本文提出EnterpriseOps-Gym,一个用于评估企业级智能体规划能力的基准。该基准包含一个容器化沙箱,内含164个数据库表和512个功能工具,以模拟真实世界的搜索摩擦。在此环境中,对14种前沿模型在8个关键业务领域(包括客户服务、人力资源和IT)的1150项专家标注任务进行评估。结果显示,表现最佳的Claude Opus 4.5仅实现37.4%的成功率。进一步分析表明,若提供人类理想计划,性能可提升14至35个百分点,暴露出战略推理是主要瓶颈。此外,智能体普遍无法拒绝不可行任务(最佳模型仅53.9%),导致意外且潜在有害的副作用。研究结论表明,当前智能体尚不具备自主企业部署的能力。更广泛而言,EnterpriseOps-Gym为提升专业工作流中智能体规划的鲁棒性提供了切实的测试平台。
原文摘要 · Abstract (English)
Large language models are shifting from passive information providers to active agents intended for complex workflows. However, their deployment as reliable AI workers in enterprise is stalled by benchmarks that fail to capture the intricacies of professional environments, specifically, the need for long-horizon planning amidst persistent state changes and strict access protocols. In this work, we introduce EnterpriseOps-Gym, a benchmark designed to evaluate agentic planning in realistic enterprise settings. Specifically, EnterpriseOps-Gym features a containerized sandbox with 164 database tables and 512 functional tools to mimic real-world search friction. Within this environment, agents are evaluated on 1,150 expert-curated tasks across eight mission-critical verticals (including Customer Service, HR, and IT). Our evaluation of 14 frontier models reveals critical limitations in state-of-the-art models: the top-performing Claude Opus 4.5 achieves only 37.4% success. Further analysis shows that providing oracle human plans improves performance by 14-35 percentage points, pinpointing strategic reasoning as the primary bottleneck. Additionally, agents frequently fail to refuse infeasible tasks (best model achieves 53.9%), leading to unintended and potentially harmful side effects. Our findings underscore that current agents are not yet ready for autonomous enterprise deployment. More broadly, EnterpriseOps-Gym provides a concrete testbed to advance the robustness of agentic planning in professional workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。