新基准让网页智能体评估更透明可靠,发现原报告数据虚高
Emergence WebVoyager: Toward Consistent and Transparent Evaluation of (Web) Agents in The Wild
- 统一任务定义、失败处理和标注规范,提升评估一致性
- 标注者一致率达95.9%,验证标准流程可靠性
- 揭露大模型在真实场景中成功率仅68.6%,远低于宣传数据
在复杂真实环境中对人工智能代理进行可靠评估,需具备鲁棒性、透明性和任务适配性的方法。本研究指出当前网页智能体评估存在任务定义模糊与操作变异等长期问题,以对WebVoyager的审计为例加以说明。为此,我们提出Emergence WebVoyager——一个增强版基准,通过明确的任务实例化、失败处理、标注与报告指南,实现标准化评估。该框架达成95.9%的标注者间一致性,显著提升任务表述与评估的清晰度与可靠性。将此框架应用于OpenAI Operator的评估,发现其在不同领域和任务类型间表现差异显著,总体成功率为68.6%,远低于OpenAI此前报告的87%,凸显本方法在实现更严格、可比的网页智能体评估中的价值。
原文摘要 · Abstract (English)
Reliable evaluation of AI agents operating in complex, real-world environments requires methodologies that are robust, transparent, and contextually aligned with the tasks agents are intended to perform. This study identifies persistent shortcomings in existing AI agent evaluation practices that are particularly acute in web agent evaluation, as exemplified by our audit of WebVoyager, including task-framing ambiguity and operational variability that hinder meaningful and reproducible performance comparisons. To address these challenges, we introduce Emergence WebVoyager, an enhanced version of the WebVoyager benchmark that standardizes evaluation methodology through clear guidelines for task instantiation, failure handling, annotation, and reporting. Emergence WebVoyager achieves an inter-annotator agreement of 95.9\%, indicating improved clarity and reliability in both task formulation and evaluation. Applying this framework to evaluate OpenAI Operator reveals substantial performance variation across domains and task types, with an overall success rate of 68.6\%, substantially lower than the 87\% previously reported by OpenAI, demonstrating the utility of our approach for more rigorous and comparable web agent evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。