系统梳理大模型代理的提示注入威胁,提出新评估基准
The Landscape of Prompt Injection Threats in LLM Agents: From Taxonomy to Analysis
- 按攻击生成策略与防御干预阶段建立分类体系
- 发现现有防御在依赖上下文的任务中普遍失效
- 新基准AgentPI揭示单一防御难兼顾安全、效率与实用
大型语言模型的发展推动了自主代理的兴起,也带来了提示注入(PI)漏洞带来的安全挑战,即恶意输入可劫持代理行为。本文通过系统文献回顾与定量分析,构建了攻击与防御的分类体系:攻击按载荷生成方式分为启发式与优化型,防御按干预阶段分为文本、模型与执行层。分析发现,多数现有防御与评测忽视了依赖运行时环境的上下文任务。为此,我们提出AgentPI基准,用于系统评估代理在上下文依赖交互下的表现。实证显示,无一防御能同时实现高可信度、高效用与低延迟;许多防御在旧基准中看似有效,实则因抑制上下文输入而无法泛化至真实场景。本文总结关键洞见与开放问题,为未来安全代理研究与部署提供指导。
原文摘要 · Abstract (English)
The evolution of Large Language Models (LLMs) has resulted in a paradigm shift towards autonomous agents, necessitating robust security against Prompt Injection (PI) vulnerabilities where untrusted inputs hijack agent behaviors. This SoK presents a comprehensive overview of the PI landscape, covering attacks, defenses, and their evaluation practices. Through a systematic literature review and quantitative analysis, we establish taxonomies that categorize PI attacks by payload generation strategies (heuristic vs. optimization) and defenses by intervention stages (text, model, and execution levels). Our analysis reveals a key limitation shared by many existing defenses and benchmarks: they largely overlook context-dependent tasks, in which agents are authorized to rely on runtime environmental observations to determine actions. To address this gap, we introduce AgentPI, a new benchmark designed to systematically evaluate agent behavior under context-dependent interaction settings. Using AgentPI, we empirically evaluate representative defenses and show that no single approach can simultaneously achieve high trustworthiness, high utility, and low latency. Moreover, we show that many defenses appear effective under existing benchmarks by suppressing contextual inputs, yet fail to generalize to realistic agent settings where context-dependent reasoning is essential. This SoK distills key takeaways and open research problems, offering structured guidance for future research and practical deployment of secure LLM agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。