arXiv:2503.07919cs.AIcs.CL2025-03被引 29

构建真实网页交互的评测基准,检验智能体处理复杂网络任务的能力。

BEARCUBS: A benchmark for computer-using web agents

  • 用111个真实网页问题测试智能体的多模态操作与信息检索能力。
  • 人类解题准确率84.7%,而最强模型仅达65.8%,差距明显。
  • 适合评估能操控键盘鼠标的真实网页智能体,推动实际应用发展。

现代网页智能体具备通过虚拟键盘和鼠标与网页交互的能力。尽管这类智能体在协助用户完成复杂任务方面潜力巨大,但如何在真实环境中评估其性能仍是重大挑战。为此,我们提出BEARCUBS——一个包含111个信息查询问题的小而强的基准,用于评估智能体在搜索、浏览和从网页中识别事实信息方面的能力。与以往基准不同,BEARCUBS要求访问实时网页内容而非合成页面,以捕捉真实网络交互的不可预测性;同时需执行多种多模态操作(如视频理解、3D导航),无法通过纯文本方式绕过。每个问题均有简短明确的答案及人工验证的浏览路径,支持透明评估智能体表现与策略。人类研究表明,这些问题可解但不简单(84.7%准确率),常见失败原因为领域知识缺失和细节遗漏。我们发现ChatGPT Agent显著优于其他计算机使用型智能体,整体准确率达65.8%(例如Operator为23.4%),展现了在真实计算机操作任务(如玩网页游戏、3D导航)上的显著进展。然而,要接近人类水平,仍需改进精细控制、复杂数据过滤与执行速度。为促进未来研究,BEARCUBS将定期更新,替换无效或污染问题,保持基准对下一代智能体的适用性。

原文摘要 · Abstract (English)

Modern web agents possess computer use abilities that allow them to interact with webpages by sending commands to a virtual keyboard and mouse. While such agents have considerable potential to assist human users with complex tasks, evaluating their capabilities in real-world settings poses a major challenge. To this end, we introduce BEARCUBS, a "smallbut mighty" benchmark of 111 information-seeking questions designed to evaluate a web agent's ability to search, browse, and identify factual information from the web. Unlike prior web agent benchmarks, solving BEARCUBS requires (1) accessing live web content rather than synthetic or simulated pages, which captures the unpredictability of real-world web interactions; and (2) performing a broad range of multimodal interactions (e.g., video understanding, 3D navigation) that cannot be bypassed via text-based workarounds. Each question in BEARCUBS has a corresponding short, unambiguous answer and a human-validated browsing trajectory, allowing for transparent evaluation of agent performance and strategies. A human study confirms that BEARCUBS questions are solvable but non-trivial (84.7% human accuracy), revealing domain knowledge gaps and overlooked details as common failure points. We find that ChatGPT Agent significantly outperforms other computer-using agents with an overall accuracy of 65.8% (compared to e.g., Operator's 23.4%), showcasing substantial progress in tasks involving real computer use, such as playing web games and navigating 3D environments. Nevertheless, closing the gap to human performance requires improvements in areas like fine control, complex data filtering, and execution speed. To facilitate future research, BEARCUBS will be updated periodically to replace invalid or contaminated questions, keeping the benchmark fresh for future generations of web agents.

网页智能体评测基准多模态交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。