评测网页智能体在安全与隐私任务上的表现,发现其自主探索能力不足。
WebSP-Eval: Evaluating Web Agents on Website Security and Privacy Tasks

- 构建200个跨28个网站的隐私安全任务数据集
- 多模型测试显示超过45%失败由开关类界面元素引发
- 适合研究智能体安全行为或网页自动化工具开发者
网页智能体可自动完成从表单填写到复杂流程(如订购生鲜)等浏览器任务。现有基准主要评估通用性能(如WebArena)或对抗恶意行为的安全性(如SafeArena),但缺乏对用户侧安全与隐私任务(如管理Cookie偏好、配置隐私设置、撤销闲置会话)的评估框架。为此,我们提出WebSP-Eval,包含:1)人工构建的200个任务实例,覆盖28个网站;2)支持账户与初始状态管理的定制化Chrome扩展代理系统;3)自动化评估器。我们使用8个基于先进多模态大模型的智能体进行评估,细粒度分析不同网站、任务类别及界面元素的表现。结果表明当前模型在自主探索方面能力有限,难以可靠完成安全隐私任务,尤其在特定类别和网站上表现差。关键发现:状态相关界面元素是主要失败原因,开关类控件导致多数模型任务失败率超45%。
原文摘要 · Abstract (English)
Web agents automate browser tasks, ranging from simple form completion to complex workflows like ordering groceries. While current benchmarks evaluate general-purpose performance~(e.g., WebArena) or safety against malicious actions~(e.g., SafeArena), no existing framework assesses an agent's ability to successfully execute user-facing website security and privacy tasks, such as managing cookie preferences, configuring privacy-sensitive account settings, or revoking inactive sessions. To address this gap, we introduce WebSP-Eval, an evaluation framework for measuring web agent performance on website security and privacy tasks. WebSP-Eval comprises 1) a manually crafted task dataset of 200 task instances across 28 websites; 2) a robust agentic system supporting account and initial state management across runs using a custom Google Chrome extension; and 3) an automated evaluator. We evaluate a total of 8 web agent instantiations using state-of-the-art multimodal large language models, conducting a fine-grained analysis across websites, task categories, and UI elements. Our evaluation reveals that current models suffer from limited autonomous exploration capabilities to reliably solve website security and privacy tasks, and struggle with specific task categories and websites. Crucially, we identify stateful UI elements are a primary reason for agent failure, with toggles causing more than 45% task failure across many models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。