arXiv:2504.18575cs.CRcs.AI2025-04NeurIPS被引 139

测试网页智能体在真实场景下对提示注入攻击的防御能力

WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks

论文配图:WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
图 1 · 摘自论文原文
  • 构建端到端评估基准WASP,模拟真实攻击场景
  • 顶级模型在86%情况下被简单提示攻击部分成功
  • 当前安全问题源于能力不足而非漏洞,适合安全研究者参考

由AI驱动的自主用户界面代理有望大幅提升人类生产力,自动完成报税、缴费等常规任务。但其代行用户操作的能力也带来了严重安全风险。现有提示注入测试或过于简化威胁场景,或赋予攻击者过强权限,或仅关注单步孤立任务。为更准确衡量安全进展,我们提出WASP——一个公开可用的端到端网页代理安全评测基准,用于评估其对抗提示注入攻击的能力。评估结果显示,即使是最先进的大模型,在高度真实的场景中仍可能被人工编写的简单提示攻击所欺骗。端到端测试揭示了一个此前未被注意的现象:尽管攻击在86%的情况下部分成功,但顶尖代理往往无法完全实现攻击目标,暴露出当前安全问题的本质是‘能力不足’而非‘设计缺陷’。

原文摘要 · Abstract (English)

Autonomous UI agents powered by AI have tremendous potential to boost human productivity by automating routine tasks such as filing taxes and paying bills. However, a major challenge in unlocking their full potential is security, which is exacerbated by the agent's ability to take action on their user's behalf. Existing tests for prompt injections in web agents either over-simplify the threat by testing unrealistic scenarios or giving the attacker too much power, or look at single-step isolated tasks. To more accurately measure progress for secure web agents, we introduce WASP -- a new publicly available benchmark for end-to-end evaluation of Web Agent Security against Prompt injection attacks. Evaluating with WASP shows that even top-tier AI models, including those with advanced reasoning capabilities, can be deceived by simple, low-effort human-written injections in very realistic scenarios. Our end-to-end evaluation reveals a previously unobserved insight: while attacks partially succeed in up to 86% of the case, even state-of-the-art agents often struggle to fully complete the attacker goals -- highlighting the current state of security by incompetence.

安全评测提示攻击网页代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。