arXiv:2409.11295cs.CRcs.AI2024-09ICLR被引 163

提出环境注入攻击,让网页代理在不知不觉中泄露用户隐私。

EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage

论文配图:EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage
图 1 · 摘自论文原文
  • 设计新型环境注入攻击,利用网页环境特性诱骗代理暴露隐私。
  • 实测显示可窃取70%特定隐私信息,16%完整请求数据。
  • 攻击隐蔽性强,现有防御难奏效,适合关注AI安全的研究者。

通用网页代理在真实网站上自主完成多种任务,显著提升效率,但涉及用户个人身份信息(PII)的网页操作可能因代理误触受损网站而引发隐私泄露,这一风险尚未被充分研究。本文首次系统探讨通用网页代理在对抗环境下的隐私风险。首先构建现实威胁模型,设定两个攻击目标:窃取用户特定PII或完整用户请求。随后提出新型攻击方法——环境注入攻击(EIA),通过生成适配代理运行环境的恶意内容实现攻击。基于Mind2Web数据集收集177个含多样PII类别的真实网页操作步骤,使用当前最先进的通用网页代理框架进行实验。结果表明,EIA在窃取特定PII时达到最高70%的攻击成功率(ASR),对完整请求的攻击成功率达16%。此外,通过分析隐蔽性并测试防御提示系统,发现该攻击难以检测和防御。值得注意的是,未适配网页的攻击可通过人工检查识别,但攻击者额外投入可使攻击无缝融入环境,使人工监督失效。因此,本文进一步讨论部署前与部署后阶段的无须人工干预的防御策略,并呼吁发展更先进的防护机制。

原文摘要 · Abstract (English)

Generalist web agents have demonstrated remarkable potential in autonomously completing a wide range of tasks on real websites, significantly boosting human productivity. However, web tasks, such as booking flights, usually involve users' PII, which may be exposed to potential privacy risks if web agents accidentally interact with compromised websites, a scenario that remains largely unexplored in the literature. In this work, we narrow this gap by conducting the first study on the privacy risks of generalist web agents in adversarial environments. First, we present a realistic threat model for attacks on the website, where we consider two adversarial targets: stealing users' specific PII or the entire user request. Then, we propose a novel attack method, termed Environmental Injection Attack (EIA). EIA injects malicious content designed to adapt well to environments where the agents operate and our work instantiates EIA specifically for privacy scenarios in web environments. We collect 177 action steps that involve diverse PII categories on realistic websites from the Mind2Web, and conduct experiments using one of the most capable generalist web agent frameworks to date. The results demonstrate that EIA achieves up to 70% ASR in stealing specific PII and 16% ASR for full user request. Additionally, by accessing the stealthiness and experimenting with a defensive system prompt, we indicate that EIA is hard to detect and mitigate. Notably, attacks that are not well adapted for a webpage can be detected via human inspection, leading to our discussion about the trade-off between security and autonomy. However, extra attackers' efforts can make EIA seamlessly adapted, rendering such supervision ineffective. Thus, we further discuss the defenses at the pre- and post-deployment stages of the websites without relying on human supervision and call for more advanced defense strategies.

隐私泄露网页代理对抗攻击AI安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。