arXiv:2601.17722cs.AI2026-01被引 4

构建企业级GUI代理的测评基准,填补专业系统任务评估空白

EntWorld: A Holistic Environment and Benchmark for Verifiable Enterprise GUI Agents

  • 基于数据库模式逆向生成真实企业流程任务
  • 1756个任务覆盖六大企业领域,成功率仅47.61%(大模型)
  • 采用SQL验证机制,确保状态转移精确可复现

多模态大语言模型使代理能在开放网络和操作系统中运行,但现有基准多聚焦消费场景(如电商、旅行预订),难以反映企业级工作流的复杂性与严谨性。企业系统面临高密度界面、严格业务逻辑约束及对精准状态一致信息检索的依赖,当前通用代理常表现不佳。为此,我们提出EntWorld,一个包含1756个任务的大规模基准,覆盖客户关系管理(CRM)、信息技术基础设施库(ITIL)和企业资源计划(ERP)等六类企业领域。不同于依赖脆弱执行轨迹或大量人工标注的数据集,EntWorld采用基于模式的任务生成框架,直接从底层数据库模式逆向推导业务逻辑,生成真实、长周期的工作流。此外,我们提出基于SQL的确定性验证机制,以严格的状态转移验证替代模糊的视觉匹配。实验表明,最先进模型(如GPT-4.1)在EntWorld上的成功率为47.61%,显著低于人类表现,凸显当前代理在企业场景中的能力鸿沟,亟需发展领域专用代理。我们公开发布EntWorld,作为下一代企业就绪数字代理研发与评估的严谨测试平台。

原文摘要 · Abstract (English)

Recent advances in Multimodal Large Language Models (MLLMs) have enabled agents to operate in open-ended web and operating system environments. However, existing benchmarks predominantly target consumer-oriented scenarios (e.g., e-commerce and travel booking), failing to capture the complexity and rigor of professional enterprise workflows. Enterprise systems pose distinct challenges, including high-density user interfaces, strict business logic constraints, and a strong reliance on precise, state-consistent information retrieval-settings in which current generalist agents often struggle. To address this gap, we introduce EntWorld, a large-scale benchmark consisting of 1,756 tasks across six representative enterprise domains, including customer relationship management (CRM), information technology infrastructure library (ITIL), and enterprise resource planning (ERP) systems. Unlike previous datasets that depend on fragile execution traces or extensive manual annotation, EntWorld adopts a schema-grounded task generation framework that directly reverse-engineers business logic from underlying database schemas, enabling the synthesis of realistic, long-horizon workflows. Moreover, we propose a SQL-based deterministic verification mechanism in building datasets that replaces ambiguous visual matching with rigorous state-transition validation. Experimental results demonstrate that state-of-the-art models (e.g., GPT-4.1) achieve 47.61% success rate on EntWorld, substantially lower than the human performance, highlighting a pronounced enterprise gap in current agentic capabilities and the necessity of developing domain-specific agents. We release EntWorld as a rigorous testbed to facilitate the development and evaluation of the next generation of enterprise-ready digital agents.

企业代理基准测试GUI智能体任务生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。