arXiv:2506.01952cs.CLcs.AI2025-06被引 28

新基准测试挑战大模型网页操作能力,专攻繁琐复杂任务

WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks

  • 设计532个真实耗时任务,聚焦记忆、计算与跨页信息追踪
  • 大模型在新任务上表现提升但仍有明显短板,验证评测有效性
  • 适合评估高阶AI代理在复杂网页操作中的真实能力

基于大语言模型的网页浏览代理能够以类人方式操作图形界面,为自动化网络任务提供透明通用框架。随着这些代理在现有基准(如WebArena)上性能快速提升,一个关键问题浮现:当前基准是否仍能准确评估日益强大的代理,尤其在更繁琐、认知负荷更高的任务中?本文提出WebChoreArena,作为WebArena的重大扩展,包含超过300小时精心设计的532项任务,明确针对劳动密集型和复杂的网页操作。它从三个关键维度系统拓展评估范围:(i) 大规模记忆,要求准确保留并检索大量观测信息;(ii) 计算,需对收集信息进行精确数学推理;(iii) 长期记忆,要求跨多个页面持续追踪信息。基于四个可复现的WebArena环境构建,确保严格兼容性,支持与先前工作的公平可控对比。实验表明,随着大模型演进,其在WebChoreArena上表现显著提升,证明该基准能更清晰地衡量前沿模型进展。然而结果也显示,即使使用GPT-5,其表现仍远低于在WebArena上的水平,凸显WebChoreArena带来的更高挑战。

原文摘要 · Abstract (English)

Powered by large language models (LLMs), web browsing agents operate graphical user interfaces in a human-like manner, offering a transparent and general framework for automating web-based tasks. As these agents rapidly improve and achieve strong performance on existing benchmarks such as WebArena, a key question arises: $\textit{Can current benchmarks still accurately evaluate the capabilities of increasingly powerful agents, especially for more tedious and cognitively demanding tasks?}$ In this paper, we present $\textbf{WebChoreArena}$, a substantial extension of WebArena designed to push beyond general browsing scenarios. WebChoreArena introduces 532 carefully curated tasks developed over 300+ hours, explicitly targeting more labor-intensive and complex web chores. It systematically expands the evaluation space along three critical dimensions: (i) $\textbf{Massive Memory}$, requiring agents to accurately retain and retrieve large amounts of information from observations; (ii) $\textbf{Calculation}$, demanding precise mathematical reasoning over collected information; and (iii) $\textbf{Long-Term Memory}$, necessitating consistent information tracking across multiple webpages. Built directly on top of the four reproducible WebArena environments, WebChoreArena ensures strict compatibility and enables fair, controlled comparisons with prior work. Our experimental results demonstrate that as LLMs evolve, significant performance improvements are observed on WebChoreArena. These findings suggest that WebChoreArena is well-suited to measure the advancement of state-of-the-art LLMs with greater clarity. Nevertheless, the results also indicate that even with GPT-5, there remains substantial room for improvement compared to WebArena, highlighting the increased challenges posed by WebChoreArena.

网页代理评测基准长程记忆大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。