arXiv:2607.06118cs.CVcs.MM2026-07

构建大规模网页代理评估基准,精准衡量导航与任务完成能力。

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation

论文配图:WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation
图 1 · 摘自论文原文
  • 引入包含800个网站的WebRetriever基准,覆盖消费、专业、企业等多领域
  • 提出NavEval框架,结合交互上下文提升评测与人工判断的一致性
  • 设计三类评估协议,全面检测代理的导航、知识交互与端到端任务能力

随着网页代理在自动化任务执行中展现日益增强的能力,构建稳健的评估框架以衡量其导航与任务完成性能已成为关键研究方向。然而,现有基准存在根本性局限:首先,规模不足且领域多样性有限,制约跨领域泛化评估;其次,主流的LLM-as-Judge方法未能充分捕捉精细的交互语义,尤其在查询构建与过滤操作方面表现不足;第三,现有基准过度关注导航成功率,忽视真实部署场景中的关键需求。为解决上述问题,我们提出WebRetriever,一个涵盖800个网站和1,550项任务的大规模基准,覆盖消费、专业及企业等多个领域,并全面覆盖用户意图模式。我们提出NavEval(导航评估)框架,利用丰富的交互上下文而非仅依赖视觉截图,实现与人类判断的最优对齐。此外,我们建立三种互补的评估协议,分别评估导航能力、知识辅助交互以及信息抽取的端到端任务完成度。大量实验分析揭示了不同评估协议间存在显著性能差异,表明仅靠导航成功无法有效预测实际应用效果。WebRetriever为代理能力提供细粒度诊断,奠定网页代理研究与开发的坚实基础。

原文摘要 · Abstract (English)

As web agents increasingly demonstrate capabilities in automated task execution, the development of robust evaluation frameworks for assessing their navigation and task completion performance has emerged as a critical research priority. However, existing benchmarks exhibit fundamental limitations. First, they suffer from insufficient scale and limited domain diversity, constraining comprehensive evaluation of cross-domain generalization. Second, prevailing LLM-as-Judge evaluation methodologies inadequately capture fine-grained interaction semantics, particularly regarding precise query formulation and filtering operations. Third, current benchmarks predominantly emphasize navigation success metrics while neglecting critical requirements for real-world deployment scenarios. To address these limitations, we introduce WebRetriever, a large-scale benchmark encompassing 800 websites and 1,550 tasks across diverse domains, including consumer, professional, and enterprise sectors, with comprehensive coverage of user intent patterns. We propose NavEval (Navigation Evaluation), a novel LLM-as-Judge framework that leverages rich interaction context beyond visual screenshots, achieving state-of-the-art alignment with human judgment across multiple evaluation datasets. Furthermore, we establish three complementary evaluation protocols that collectively provide holistic assessment of web agent capabilities: navigation proficiency, knowledge-assisted interaction, and end-to-end task completion with information extraction. Extensive experimental analysis reveals substantial performance disparities across evaluation protocols, demonstrating that navigation success alone is an insufficient predictor of real-world application effectiveness. WebRetriever delivers fine-grained diagnostic insights into agent capabilities and establishes a rigorous foundation for advancing web agent research and development.

网页代理评估基准LLM评估任务完成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。