arXiv:2605.17634cs.CRcs.CL2026-05被引 8

AI代理永远可能中招提示注入攻击,因当前防御机制有根本性缺陷。

AI Agents May Always Fall for Prompt Injections

  • 用上下文完整性理论重新理解提示注入攻击
  • 三种新攻击场景让代理违背信息流规范
  • 提出可指导未来智能体对齐设计的理论框架

提示注入是部署中AI代理最严重的漏洞。尽管近期已有进展,我们发现主流防御范式(数据-指令分离)既无法检测通过上下文操纵实施的攻击,又会损害上下文相关的正常行为。我们从上下文完整性(CI)这一隐私理论视角重新审视提示注入,该理论以符合上下文规范来判断信息流动是否合规。这解释了现有防御试图修补的攻击类型,并预测未来代理将面临的高级攻击。我们设计了独特的良性与攻击场景,迫使代理违反规范:(1)歪曲信息流,(2)操纵规范本身,或(3)混合多重信息流。这一重构揭示了一个不可能性结果:攻击者总能构造出使被拦截流看似合法的上下文;而防御者若收紧规范,则会误拦真正合法的流。研究显示,当前研究仅覆盖未来攻击面的一小部分。相反,通过CI,我们提供了一个评估上下文敏感失效的系统性框架,并为前沿自主代理设计具备CI意识的对齐方案。

原文摘要 · Abstract (English)

Prompt injection is the most critical vulnerability in deployed AI agents. Despite recent progress, we show that the prevailing defense paradigm (data-instruction separation) both fails to detect attacks that operate through contextual manipulation and degrades contextually appropriate behavior. We then recast prompt injection via the lens of Contextual Integrity (CI), a privacy theory that judges information flow compliance with contextual norms. This explains types of attacks that current defenses attempt to patch and predict advanced ones future agents will face. We develop unique benign and attack scenarios that force an agent to violate the norms by (1) misrepresenting the flow, (2) manipulating norms, or (3) mixing multiple flows. This reframing suggests an impossibility result: an adversary can always construct a context under which a blocked flow appears legitimate, or a defender who tightens norms will block genuinely legitimate flows. Our findings suggest that current research addresses a shrinking fraction of future attack surfaces. Instead, through CI, we offer a principled framework for evaluating context-sensitive failures, and designing CI-aware alignment for the frontier autonomous agents.

提示注入安全漏洞上下文完整性智能体对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。