评测跨应用办公的自动化界面智能体,发现当前模型表现远不及人类。
WindowsWorld: A Process-Centric Benchmark of Autonomous GUI Agents in Professional Cross-Application Environments

- 用16种职业设计多应用复杂任务,模拟真实工作流程。
- 多应用任务成功率低于21%,且在3个以上应用间推理时几乎失败。
- 适合评估智能体在真实办公场景中的协作与决策能力。
尽管图形界面(GUI)智能体在常见计算机任务中表现优异,但现有基准大多聚焦于孤立的单应用任务,忽略了专业工作中跨多个应用协同完成复杂流程的关键需求。为填补这一空白,我们提出名为WindowsWorld的跨应用工作流基准,系统评估智能体在模拟真实职业活动的多步骤任务中的表现。该方法基于16种职业构建四类难度的任务,包含中间检查点,经人工审核后在仿真环境中执行。最终基准包含181项任务,平均每个任务有5.0个子目标,覆盖17个常用桌面应用,其中78%的任务天然涉及多应用协作。实验结果表明:1)所有主流大模型和智能体在多应用任务上的成功率均低于21%,远低于单应用任务表现;2)在涉及≥3个应用的条件判断与推理任务中,智能体普遍在早期子目标处停滞;3)执行效率低下,任务常因远超人类操作步数而失败。代码、数据与评估资源已开源。
原文摘要 · Abstract (English)
While GUI agents have shown impressive capabilities in common computer-use tasks such as OSWorld, current benchmarks mainly focus on isolated and single-application tasks. This overlooks a critical real-world requirement of coordinating across multiple applications to accomplish complex profession-specific workflows. To bridge this gap, we present a computer-use benchmark in cross-application workflows, named WindowsWorld, designed to systematically assess GUI Agents on complex multi-step tasks that mirror real-world professional activities. Our methodology uses a multi-agent framework steered by 16 occupations to generate four difficulty-level tasks with intermediate inspection, which are then refined by human review and executed in a simulated environment. The resulting benchmark contains 181 tasks with an average of 5.0 sub-goals across 17 common desktop applications, of which 78% are inherently multi-application. Experimental results of leading large models and agents show that: 1) All computer-use agents perform poorly on multi-application tasks (< 21% success rate), far below the performance of simple single-app tasks; 2) They largely fail at tasks requiring conditional judgment and reasoning across $\geq$ 3 applications, stalling at early sub-goals; 3) Low execution efficiency, where tasks often fail despite far exceeding human step limits. Code, benchmark data, and evaluation resources are available at github.com/HITsz-TMG/WindowsWorld.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。