首次系统分析大模型代理的间接提示攻击防御框架,揭示其漏洞并提出新攻击方法。
Taxonomy, Evaluation and Exploitation of IPI-Centric LLM Agent Defense Frameworks
- 构建五维分类体系,梳理现有防御方法
- 发现六类防御失效根源,攻击成功率显著提升
- 适合安全研究者和大模型开发者参考
具备函数调用能力的大语言模型(LLM)代理正被广泛部署,但易受间接提示攻击(IPI)的影响,导致工具调用被劫持。针对此问题,涌现出多种以IPI为中心的防御框架,但缺乏统一分类与全面评估。本文通过知识体系化(SoK)研究,首次系统分析了这些防御框架。我们提出一个涵盖五个维度的综合分类体系,并对代表性防御方案进行了安全性和可用性评估。通过对防御失败案例的分析,识别出六种导致防御被绕过的根本原因。基于此,我们设计了三种新型自适应攻击,在特定框架上显著提升了攻击成功率,凸显了现有防御机制的严重缺陷。本研究为未来更安全、更易用的IPI中心防御框架发展奠定了基础并提供了关键洞见。
原文摘要 · Abstract (English)
Large Language Model (LLM)-based agents with function-calling capabilities are increasingly deployed, but remain vulnerable to Indirect Prompt Injection (IPI) attacks that hijack their tool calls. In response, numerous IPI-centric defense frameworks have emerged. However, these defenses are fragmented, lacking a unified taxonomy and comprehensive evaluation. In this Systematization of Knowledge (SoK), we present the first comprehensive analysis of IPI-centric defense frameworks. We introduce a comprehensive taxonomy of these defenses, classifying them along five dimensions. We then thoroughly assess the security and usability of representative defense frameworks. Through analysis of defensive failures in the assessment, we identify six root causes of defense circumvention. Based on these findings, we design three novel adaptive attacks that significantly improve attack success rates targeting specific frameworks, demonstrating the severity of the flaws in these defenses. Our paper provides a foundation and critical insights for the future development of more secure and usable IPI-centric agent defense frameworks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。