arXiv:2608.08392cs.AIcs.CL2026-08中稿 · COLM

构建跨网站浏览器代理评估基准,测试复杂操作与视觉理解能力

CAP: A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception

论文配图:CAP: A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception
图 1 · 摘自论文原文
  • 将网站抽象为结构化卡片,重组为真实跨站任务流
  • 420个任务覆盖108个真实网站,验证现有代理成功率低
  • 聚焦视觉感知与复杂交互,适合研究智能代理的学者

大型语言模型正被用作通过浏览器与网络交互的自主代理。尽管近期进展依赖于评估端到端任务成功率的基准,但这些评估大多忽略了真实网页浏览中的两大核心难点:对丰富用户界面的复杂操作,以及对动态渲染内容的视觉感知,尤其在跨网站流程中。我们提出CAP,一个可扩展的基准,用于评估浏览器代理在需非平凡用户界面交互和视觉理解的跨站点、类人任务上的表现。具体地,采用分解-重构流程,先将每个网站抽象为结构化站点卡,包含用户功能、复杂执行操作及感知需求,再将这些组件重组为真实跨站工作流。每项任务因此建立在各网站的具体操作基础上,实现细粒度诊断。基于此框架,我们在108个真实网站、24个领域内构建了420个任务,并经过严格质量控制。使用可验证的代理作为裁判评估框架对最先进浏览器代理进行实验,结果显示成功率低下,表明以感知为主的交互仍是主要瓶颈,揭示当前代理与真实网络浏览需求之间存在显著差距。

原文摘要 · Abstract (English)

Large language models are increasingly deployed as autonomous agents that interact with the web through browsers. While recent progress has been driven by benchmarks that evaluate end-to-end task success, these evaluations largely overlook two fundamental sources of difficulty in real web browsing: complex actions over rich user interfaces and visual perception of dynamically rendered content, especially in workflows that span multiple websites. We introduce CAP, a scalable benchmark for evaluating browser agents on cross-site, human-like web tasks that require non-trivial UI interactions and visual understanding. Specifically, we adopt a decomposition-and-recomposition pipeline that first abstracts each website into a structured site card capturing user-facing functions, complex execution operations, and perceptual requirements, and then recomposes these components into realistic cross-site workflows. Each task is therefore grounded in multiple specific operations on each website, enabling fine-grained diagnosis. Built on this framework, we construct 420 tasks across 108 real-world websites and 24 domains under careful quality control. Experiments on state-of-the-art browser agents using our verifiable agent-as-a-judge evaluation framework show low success rates and reveal that perception-heavy interactions remain a major bottleneck, exposing substantial gaps between current agents and real-world web browsing demands.

浏览器代理跨站任务视觉理解评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。