评测智能体在真实桌面环境中跨界面协同完成复杂任务的能力
WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces

- 构建包含114个真实任务的混合接口基准,覆盖8大工作场景
- 顶尖模型在长周期任务中通过率仅41.2%,表明能力仍有巨大差距
- 新增轨迹感知评估器,可识别伪造证据和硬编码行为,避免结果虚高
计算机使用智能体(CUAs)正运行在结合图形界面、命令行、代码编辑、浏览器及外部工具的混合环境中。现有基准多将这些接口视为独立能力进行评估,导致长周期跨界面协同任务缺乏有效测试。为此,我们提出WeaveBench,一个面向真实世界、具有长期跨度的混合接口基准,包含114项任务,覆盖8个实际工作领域,基于真实用户请求并具备公开可验证成果。每项任务要求智能体在同一轨迹中整合GUI观测/操作与CLI/代码执行。我们在部署的CLI智能体运行时中,于真实Ubuntu桌面环境上进行评估,并引入最小化桌面控制插件。同时提出配套的轨迹感知评估器,检查交付成果、文件、截图、日志及操作轨迹,能检测伪造视觉证据或硬编码指标等捷径行为。在前沿模型-运行时组合中,最高通过率为41.2%,显示该基准尚未饱和。轨迹感知评估器进一步表明,仅以结果评分会严重夸大智能体表现。总体而言,WeaveBench揭示了当前CUA评估中的关键缺口,并提供了一个有效平台,用于衡量智能体在长周期真实任务中协调GUI、CLI与代码操作的能力。
原文摘要 · Abstract (English)
Computer-use agents (CUAs) increasingly operate in runtimes that combine visual desktop control, command-line execution, code editing, browsers, and external tools. Existing benchmarks, however, often evaluate these interfaces as separable capabilities, leaving long-horizon cross-interface orchestration under-tested. Thus, we introduce WeaveBench, a long-horizon hybrid-interface benchmark with 114 tasks across 8 real-world work domains, grounded in real user requests and publicly verifiable artifacts. Each task requires agents to combine GUI observations/actions with CLI/code operations within a single trajectory. We evaluate these tasks on a real Ubuntu desktop inside deployed CLI-agent runtimes, augmented with a minimal desktop-control plugin. We also propose a companion trajectory-aware judge that inspects deliverables, files, screenshots, logs, and action traces, while detecting shortcut behaviors such as fabricated visual evidence or hard-coded metrics. Across frontier model-runtime pairings, the best PassRate reaches only 41.2%, showing the benchmark remains far from saturated. The trajectory-aware judge further reveals that outcome-only grading substantially overestimates agent performance. Overall, WeaveBench exposes a critical gap in CUA evaluation and provides an effective testbed to measure whether agents can orchestrate GUI, CLI, and code operations across long-horizon real-world tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。