arXiv:2502.18356cs.LG2025-02被引 9

构建首个综合性网页浏览AI评估基准,揭示当前模型与人类差距。

WebGames: Challenging General-Purpose Web-Browsing AI Agents

  • 设计50+交互挑战,覆盖基础操作到复杂任务
  • 顶尖AI成功率仅43.1%,远低于人类95.7%
  • 开源轻量框架,适合研究者快速测试新模型

我们提出WebGames,一个涵盖50多个交互式挑战的综合性基准,用于评估通用网页浏览AI代理。这些挑战对人类简单,但系统性地检验当前AI在基础浏览器操作、高级输入处理、认知任务、工作流自动化及互动娱乐等方面的局限性。该框架通过隔离环境消除外部依赖,确保可复现的评估和可验证的真实答案。我们评估了GPT-4o、Claude Computer-Use、Gemini-1.5-Pro和Qwen2-VL等主流视觉语言模型,结果显示最佳AI系统仅达43.1%成功率,远低于人类95.7%的表现,暴露出当前AI在处理人类习以为常的网页交互模式上的根本缺陷。基准已公开于webgames.convergence.ai,采用轻量级客户端实现,支持快速评估迭代。其模块化架构与标准化挑战规范,为提升网页浏览代理能力提供了可靠评估基础。

原文摘要 · Abstract (English)

We introduce WebGames, a comprehensive benchmark suite designed to evaluate general-purpose web-browsing AI agents through a collection of 50+ interactive challenges. These challenges are specifically crafted to be straightforward for humans while systematically testing the limitations of current AI systems across fundamental browser interactions, advanced input processing, cognitive tasks, workflow automation, and interactive entertainment. Our framework eliminates external dependencies through a hermetic testing environment, ensuring reproducible evaluation with verifiable ground-truth solutions. We evaluate leading vision-language models including GPT-4o, Claude Computer-Use, Gemini-1.5-Pro, and Qwen2-VL against human performance. Results reveal a substantial capability gap, with the best AI system achieving only 43.1% success rate compared to human performance of 95.7%, highlighting fundamental limitations in current AI systems' ability to handle common web interaction patterns that humans find intuitive. The benchmark is publicly available at webgames.convergence.ai, offering a lightweight, client-side implementation that facilitates rapid evaluation cycles. Through its modular architecture and standardized challenge specifications, WebGames provides a robust foundation for measuring progress in development of more capable web-browsing agents.

网页浏览AI评估基准测试多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。