评测AI在真实专业软件中完成长流程任务的能力,发现当前模型成功率不足35%。
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields

- 构建面向专业领域长流程任务的GUI基准测试集Workflow-GYM
- 顶尖模型在真实工作流上成功率仅略超30%,普遍存在流程中断和目标漂移
- 揭示了当前AI代理在专业软件环境中的长期一致性短板,适合研究人参考
近年来,AI代理在处理日益复杂的现实任务方面迅速发展。然而,现有评估基准很少检验代理是否能操作图形用户界面,以完成跨多个领域的长周期、高价值专业工作流。当前的GUI基准仍主要聚焦于通用软件、相对简单的应用和短周期任务,尚不清楚现代代理能否根据用户指令自主操作特定领域专业软件并以端到端方式完成经济价值高的工作。为填补这一空白,我们提出Workflow-GYM,一个聚焦于专业领域与专用软件环境的长周期GUI任务基准。对最先进模型的大量实验表明,即使最强模型的成功率也仅略高于30%,凸显当前专业长周期GUI工作流对代理而言仍极具挑战性。进一步分析显示,当前代理难以保持长周期工作流的一致性,常出现阶段遗漏、错误传播、目标漂移以及对专业软件环境理解不足等问题。研究结果为当前代理系统局限性提供了重要洞察,并指明下一代GUI代理研究的关键方向。
原文摘要 · Abstract (English)
Recent years have witnessed the rapid evolution of AI agents toward handling increasingly complex, real-world tasks. However, existing benchmarks rarely evaluate whether agents can operate graphical user interfaces to complete long-horizon, high-value professional workflows across diverse domains. Current GUI benchmarks still predominantly focus on general-purpose software, relatively simple applications, and short-horizon tasks, leaving it largely unknown whether modern agents can follow user instructions to autonomously operate domain-specific professional software and accomplish economically valuable work in an end-to-end manner. To bridge this gap, we introduce Workflow-GYM, a benchmark for long-horizon GUI tasks centered on professional domains and specialized software environments. Through extensive experiments on state-of-the-art models, we find that even the strongest models achieve only slightly above 30% success rates, highlighting that professional long-horizon GUI workflows remain highly challenging for current GUI agents. Further analysis reveals that current agents struggle to maintain long-horizon workflow consistency, frequently exhibiting workflow stage omission, error propagation, objective drift, and insufficient understanding of professional software environments. Our findings provide important insights into the limitations of current agent systems and suggest key directions for the next generation of GUI-agent research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。