测试企业级自动化代理在复杂任务中的可靠性,发现现有模型表现断崖式下降。
UI-CUBE: Enterprise-Grade Computer Use Agent Benchmarking Beyond Task Accuracy to Operational Reliability
- 构建226个任务的基准,覆盖简单操作与企业级复杂流程。
- 简单任务成功率67-85%,复杂流程骤降至9-19%,远低于人类表现。
- 揭示内存管理与状态协调等架构缺陷,适合评估生产可用性。
当前计算机使用代理(CUA)基准主要衡量任务完成度,但难以评估企业部署就绪性,忽视生产系统所需的运行可靠性。我们提出UI-CUBE(UiPath Computer Use BEnchmark),一个包含226个任务的系统性基准,分为两个难度层级,用于暴露现有CUA的基础架构局限。评估涵盖136个简单界面交互、50个复制粘贴任务及40个企业应用场景,覆盖系统性界面变化、多分辨率测试,并通过应用状态自动验证任务成功。对五种前沿模型的评估显示,性能呈现陡峭悬崖式下降而非渐进退化:简单任务成功率为67-85%(人类为97.9%),复杂工作流则骤降至9-19%。无经验人类评估者在复杂任务上仅达61.2%,尽管简单任务接近完美。这一非连续表现模式——简单任务达人类68-87%表现,复杂流程仅15-32%——表明根本性架构缺陷存在于记忆管理、分层规划与状态协调,而非可通过训练或提示优化的渐进差距。UI-CUBE可作为企业就绪性诊断工具,揭示当前CUA虽能操控单个界面元素,却尚无法作为可靠的工作流自动化工具。这些发现为开发可投入生产的复杂企业流程自动化代理提供关键架构洞见。
原文摘要 · Abstract (English)
While current Computer Use Agent (CUA) benchmarks measure task completion effectively, they provide limited assessment of enterprise deployment readiness, emphasizing functional correctness over the operational reliability required for production systems. We present UI-CUBE (UiPath Computer Use BEnchmark), a systematic benchmark comprising 226 tasks across two difficulty tiers designed to expose fundamental architectural limitations in current CUAs. Our evaluation covers simple UI interactions (136 tasks) and complex workflows including copy-paste tasks (50 tasks) and enterprise application scenarios (40 tasks), with systematic interface variation coverage, multi-resolution testing and automated validation of task success through the application state. Evaluation of five state-of-the-art models reveals a sharp capability cliff rather than gradual performance degradation. Simple UI interactions achieve 67-85% success rates (compared to 97.9% human performance), but complex workflows drop precipitously to 9-19%. Human evaluators with no prior application experience achieve only 61.2% on complex tasks despite near-perfect performance on simple tasks, establishing realistic performance ceilings. This discontinuous performance pattern -- where agents achieve 68-87% of human performance on simple tasks but only 15-32% on complex workflows -- indicates fundamental architectural limitations in memory management, hierarchical planning, and state coordination rather than incremental capability gaps addressable through better training or prompting. UI-CUBE functions as an enterprise-readiness diagnostic, revealing that while current CUAs can manipulate individual interface elements, they cannot yet function as reliable workflow automation tools. These findings provide architectural insights essential for developing production-ready CUAs capable of managing complex, multi-step enterprise processes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。