现有评测体系常误判计算机操作智能体表现,15.3%失败结果为错误判定。
How Benchmarks Mis-Score Computer-Use Agents
- 构建评测可靠性框架,覆盖任务设计到报告全流程
- 审计150条轨迹发现15.3%失败判定错误,含任务缺陷与评估误判
- 提出分阶段评估规则,适用于长周期智能体评测
计算机使用智能体(CUA)正被部署于网页浏览和桌面软件操作,但其基准评分仍依赖脆弱的脚本化评判机制。评分过程存在任务过时、轨迹遗漏关键视觉证据、评估器拒绝有效方案以及汇总报告掩盖失败原因等问题。本文将这些缺陷归纳为涵盖任务构建、轨迹观测、评分与报告的可靠性框架。我们审计了来自五个网页、企业工作流及桌面控制基准的150条公开失败轨迹,发现15.3%的失败判定错误:其中10.7%为评估器假阴性,4.7%源于任务本身缺陷。对于真实失败,三层次诊断分类显示,验证/反馈与规划失误远超执行/接地错误;单一成功率指标无法充分解释失败模式。研究将结论延伸至新型长周期CUA基准,并提出针对性评估设计准则。
原文摘要 · Abstract (English)
Computer-use agents (CUA) are being deployed to browse the web and operate desktop software, yet their benchmark scores are still commonly produced by brittle scripted oracles. A score is the output of a pipeline in which tasks can be stale, trajectories can omit decisive visual evidence, evaluators can reject valid alternatives, and aggregate reports can hide the cause of failure. We organize these problems into a reliability framework spanning task construction, trajectory observation, scoring, and reporting. We then audit 150 public failure-scored trajectories from five web, enterprise-workflow, and desktop-control benchmarks, find that 15.3\% of FAIL verdicts are wrong: 10.7\% are evaluator false negatives and 4.7\% are broken tasks. For genuine failures, a three-tier diagnostic taxonomy shows that verification/feedback and planning failures dominate execution/grounding errors, while a single scalar success rate can not explain. We connect these findings to newer long-horizon CUA benchmarks and derive stage-specific design rules for CUA evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。