arXiv:2511.20597cs.LGcs.AI2025-11被引 26

提出真实场景下的提示注入攻击基准,评估并设计多层防御策略。

BrowseSafe: Understanding and Preventing Prompt Injection Within AI Browser Agents

  • 构建含复杂干扰的网页级提示注入攻击基准
  • 验证主流模型在真实攻击下防御失效,平均成功率超70%
  • 提出架构+模型联合的纵深防御方案,适合安全研究者与开发团队

将AI代理集成至网络浏览器带来了超越传统网络应用威胁模型的安全挑战。先前研究已识别出提示注入作为新型攻击向量,但其在真实环境中的影响仍不明确。本文分析了提示注入攻击现状,构建了一个嵌入真实HTML载荷的攻击基准。该基准超越以往工作,强调能引发实际行为而非仅文本输出的注入,并以与真实世界代理所遇相似的复杂度和干扰频率呈现攻击载荷。利用此基准,我们对现有防御措施进行了全面实证评估,测试其在一系列前沿AI模型上的表现。提出一种包含架构与模型级防御的多层次防护策略,以应对持续演化的提示注入攻击。本工作为通过纵深防御设计实用、安全的网络代理提供了蓝图。

原文摘要 · Abstract (English)

The integration of artificial intelligence (AI) agents into web browsers introduces security challenges that go beyond traditional web application threat models. Prior work has identified prompt injection as a new attack vector for web agents, yet the resulting impact within real-world environments remains insufficiently understood. In this work, we examine the landscape of prompt injection attacks and synthesize a benchmark of attacks embedded in realistic HTML payloads. Our benchmark goes beyond prior work by emphasizing injections that can influence real-world actions rather than mere text outputs, and by presenting attack payloads with complexity and distractor frequency similar to what real-world agents encounter. We leverage this benchmark to conduct a comprehensive empirical evaluation of existing defenses, assessing their effectiveness across a suite of frontier AI models. We propose a multi-layered defense strategy comprising both architectural and model-based defenses to protect against evolving prompt injection attacks. Our work offers a blueprint for designing practical, secure web agents through a defense-in-depth approach.

提示注入AI安全浏览器代理防御策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。