arXiv:2507.00938cs.IRcs.AI2025-07

构建静态网页任务基准,评估大模型网络代理的稳定表现

WebArXiv: Evaluating Multimodal Agents on Time-Invariant arXiv Tasks

  • 用arXiv固定快照构建275个不变任务,确保评估可复现
  • 发现代理过度依赖历史记录导致失败,性能差异显著
  • 提出轻量动态反思机制,提升决策时对历史的灵活调用

大语言模型的发展催生了能自主操作真实网站的网络代理,但现有评估基准因内容动态性或模拟简化而不可靠。本文提出WebArXiv,一个基于arXiv平台的静态、时间不变基准,包含275个网页任务。通过固定网页快照与确定性真值,实现可复现且可靠的评估。行为分析揭示常见失败模式:僵化历史反射,即代理过度依赖固定交互历史。为此,我们提出轻量级动态反思机制,使代理在决策时可选择性检索相关历史步骤。在WebArXiv上评估了10个顶尖网络代理,结果表明各代理表现差异明显,且所提反思策略有效提升性能。

原文摘要 · Abstract (English)

Recent progress in large language models (LLMs) has enabled the development of autonomous web agents capable of navigating and interacting with real websites. However, evaluating such agents remains challenging due to the instability and inconsistency of existing benchmarks, which often rely on dynamic content or oversimplified simulations. In this work, we introduce WebArXiv, a static and time-invariant benchmark comprising 275 web-based tasks grounded in the arXiv platform. WebArXiv ensures reproducible and reliable evaluation by anchoring tasks in fixed web snapshots with deterministic ground truths and standardized action trajectories. Through behavioral analysis, we identify a common failure mode, Rigid History Reflection, where agents over-rely on fixed interaction histories. To address this, we propose a lightweight dynamic reflection mechanism that allows agents to selectively retrieve relevant past steps during decision-making. We evaluate ten state-of-the-art web agents on WebArXiv. Results demonstrate clear performance differences across agents and validate the effectiveness of our proposed reflection strategy.

网络代理评估基准LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。