评估现有网页代理真实能力,发现此前结果过于乐观。
An Illusion of Progress? Assessing the Current State of Web Agents
- 构建在线动态基准Online-Mind2Web,覆盖300个真实任务
- 新评测方法与人工判断一致率达85%,显著优于旧方法
- 首次全面对比现役网页代理,揭示其真实优劣
随着数字化和云技术的发展,网络在现代社会中的重要性日益提升。基于大语言模型的自主网页代理在工作自动化方面具有巨大潜力,因此准确衡量其能力进展至关重要。本文对当前网页代理的能力状态进行了全面而严谨的评估,结果揭示了与以往报告相比更为保守的现实能力图景,表明先前研究存在过度乐观。这一差距源于现有基准的不足。为此,我们提出了Online-Mind2Web,一个包含300个多样化、真实场景任务的在线评估基准,覆盖136个网站,可模拟真实用户使用代理的环境。为支持更高效的评估与开发,我们还设计了一种新的LLM-as-a-Judge自动评估方法,其与人工判断的一致率可达约85%,远高于现有方法。最后,我们首次系统性地对比分析了当前主流网页代理,明确指出其优势与局限,以推动未来研究。
原文摘要 · Abstract (English)
As digitalization and cloud technologies evolve, the web is becoming increasingly important in the modern society. Autonomous web agents based on large language models (LLMs) hold a great potential in work automation. It is therefore important to accurately measure and monitor the progression of their capabilities. In this work, we conduct a comprehensive and rigorous assessment of the current state of web agents. Our results depict a very different picture of the competency of current agents, suggesting over-optimism in previously reported results. This gap can be attributed to shortcomings in existing benchmarks. We introduce Online-Mind2Web, an online evaluation benchmark consisting of 300 diverse and realistic tasks spanning 136 websites. It enables us to evaluate web agents under a setting that approximates how real users use these agents. To facilitate more scalable evaluation and development, we also develop a novel LLM-as-a-Judge automatic evaluation method and show that it can achieve around 85% agreement with human judgment, substantially higher than existing methods. Finally, we present the first comprehensive comparative analysis of current web agents, highlighting both their strengths and limitations to inspire future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。