构建统一网页智能体评估生态,实现跨基准可比测试。
The BrowserGym Ecosystem for Web Agent Research
- 基于BrowserGym框架统一多个网页智能体基准测试
- 6大主流模型在6个基准上对比,Claude-3.5-Sonnet综合领先
- 适合研究网页自动化与大模型应用的开发者与学者
BrowserGym生态系统应对网页智能体评估与基准测试日益增长的需求,尤其针对依赖自动化和大语言模型(LLMs)的智能体。现有基准存在碎片化与评价方法不一致问题,难以实现可靠比较与可复现结果。此前工作(Drouin et al., 2024)提出BrowserGym,通过定义清晰的观测与动作空间,提供类gym的统一环境。本文扩展该生态,整合文献中多个基准,并引入AgentLab框架,支持智能体创建、测试与分析。新生态兼顾灵活性与一致性,支持新基准接入与全面实验管理。作为验证,我们首次开展大规模多基准实验,对比6个前沿LLMs在6个公开网页智能体基准上的表现。结果显示,Claude-3.5-Sonnet在几乎所有任务中领先,仅在视觉相关任务上低于GPT-4o。尽管进步显著,当前模型在真实复杂网页环境中仍面临挑战,表明构建鲁棒高效网页智能体难度较大。
原文摘要 · Abstract (English)
The BrowserGym ecosystem addresses the growing need for efficient evaluation and benchmarking of web agents, particularly those leveraging automation and Large Language Models (LLMs). Many existing benchmarks suffer from fragmentation and inconsistent evaluation methodologies, making it challenging to achieve reliable comparisons and reproducible results. In an earlier work, Drouin et al. (2024) introduced BrowserGym which aims to solve this by providing a unified, gym-like environment with well-defined observation and action spaces, facilitating standardized evaluation across diverse benchmarks. We propose an extended BrowserGym-based ecosystem for web agent research, which unifies existing benchmarks from the literature and includes AgentLab, a complementary framework that aids in agent creation, testing, and analysis. Our proposed ecosystem offers flexibility for integrating new benchmarks while ensuring consistent evaluation and comprehensive experiment management. As a supporting evidence, we conduct the first large-scale, multi-benchmark web agent experiment and compare the performance of 6 state-of-the-art LLMs across 6 popular web agent benchmarks made available in BrowserGym. Among other findings, our results highlight a large discrepancy between OpenAI and Anthropic's latests models, with Claude-3.5-Sonnet leading the way on almost all benchmarks, except on vision-related tasks where GPT-4o is superior. Despite these advancements, our results emphasize that building robust and efficient web agents remains a significant challenge, due to the inherent complexity of real-world web environments and the limitations of current models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。