首个全面评估视觉语言模型网页代理安全性的基准,覆盖六类攻击场景。
SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents
- 构建六个真实模拟网页环境,涵盖电商、论坛等场景
- 测试9个主流模型均暴露于细微对抗攻击,安全与专用性存在权衡
- 从推理、行为到任务结果多层分析,实现细粒度风险评估
基于大视觉语言模型(LVLM)的网页代理正成为自动化复杂在线任务的强大工具。然而在真实环境中部署时面临严重安全风险,亟需建立评估基准。现有基准覆盖不全,多局限于用户级提示操控等狭窄场景,难以捕捉代理漏洞的全貌。为此,我们提出首个面向LVLM网页代理的综合性安全评估基准 ool{}。该基准包含六个模拟但真实的网页环境(如电商平台、社区论坛),涵盖2,970条高质量轨迹,覆盖多样化任务与攻击设置,并定义了六类攻击向量,涵盖用户级与环境级操纵。此外,引入多层级评估协议,从内部推理、行为轨迹和任务结果三个维度分析代理失败,实现超越简单成功指标的细粒度风险分析。我们在9个代表性LVLM上开展大规模实验,涵盖通用型、代理专用型和GUI引导型三类模型。结果表明,所有被测代理均对细微对抗操纵高度敏感,揭示出模型专业化程度与安全性之间的关键权衡。 ool{}通过提供全面的基准套件与多层评估流程,以及对现代LVLM网页代理安全挑战的实证洞察,为可信网页代理部署奠定了基础。
原文摘要 · Abstract (English)
Large vision-language model (LVLM)-based web agents are emerging as powerful tools for automating complex online tasks. However, when deployed in real-world environments, they face serious security risks, motivating the design of security evaluation benchmarks. Existing benchmarks provide only partial coverage, typically restricted to narrow scenarios such as user-level prompt manipulation, and thus fail to capture the broad range of agent vulnerabilities. To address this gap, we present \tool{}, the first holistic benchmark for evaluating the security of LVLM-based web agents. \tool{} first introduces a unified evaluation suite comprising six simulated but realistic web environments (\eg, e-commerce platforms, community forums) and includes 2,970 high-quality trajectories spanning diverse tasks and attack settings. The suite defines a structured taxonomy of six attack vectors spanning both user-level and environment-level manipulations. In addition, we introduce a multi-layered evaluation protocol that analyzes agent failures across three critical dimensions: internal reasoning, behavioral trajectory, and task outcome, facilitating a fine-grained risk analysis that goes far beyond simple success metrics. Using this benchmark, we conduct large-scale experiments on 9 representative LVLMs, which fall into three categories: general-purpose, agent-specialized, and GUI-grounded. Our results show that all tested agents are consistently vulnerable to subtle adversarial manipulations and reveal critical trade-offs between model specialization and security. By providing (1) a comprehensive benchmark suite with diverse environments and a multi-layered evaluation pipeline, and (2) empirical insights into the security challenges of modern LVLM-based web agents, \tool{} establishes a foundation for advancing trustworthy web agent deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。