系统梳理88篇论文,构建更完整的提示注入防御分类体系。
A Systematic Literature Review on LLM Defenses Against Prompt Injection and Jailbreaking: Expanding NIST Taxonomy
- 基于NIST框架扩展防御类别,覆盖更多新方法。
- 汇总88项研究在具体模型和数据集上的防御效果。
- 提供开源、通用的防御方案清单,适合开发者落地使用。
生成式人工智能与大语言模型的快速发展带来了提示注入攻击等新型安全威胁,恶意输入可能导致数据泄露或非法操作。为应对快速演进的攻防技术,本文首次系统性回顾了88项提示注入防御研究。在借鉴NIST对抗机器学习报告的基础上,本文拓展了防御分类体系,纳入更多未被现有综述涵盖的方法;提出统一术语与分类标准,促进研究一致性;并构建完整防御方案目录,记录各方案在特定大语言模型和攻击数据集上的定量表现,同时标注其是否开源及是否模型无关。该目录与指南旨在为研究人员推进对抗机器学习领域,以及开发者在生产环境中部署有效防御提供实用参考。
原文摘要 · Abstract (English)
The rapid advancement and widespread adoption of generative artificial intelligence (GenAI) and large language models (LLMs) has been accompanied by the emergence of new security vulnerabilities and challenges, such as jailbreaking and other prompt injection attacks. These maliciously crafted inputs can exploit LLMs, causing data leaks, unauthorized actions, or compromised outputs, for instance. As both offensive and defensive prompt injection techniques evolve quickly, a structured understanding of mitigation strategies becomes increasingly important. To address that, this work presents the first systematic literature review on prompt injection mitigation strategies, comprehending 88 studies. Building upon NIST's report on adversarial machine learning, this work contributes to the field through several avenues. First, it identifies studies beyond those documented in NIST's report and other academic reviews and surveys. Second, we propose an extension to NIST taxonomy by introducing additional categories of defenses. Third, by adopting NIST's established terminology and taxonomy as a foundation, we promote consistency and enable future researchers to build upon the standardized taxonomy proposed in this work. Finally, we provide a comprehensive catalog of the reviewed prompt injection defenses, documenting their reported quantitative effectiveness across specific LLMs and attack datasets, while also indicating which solutions are open-source and model-agnostic. This catalog, together with the guidelines presented herein, aims to serve as a practical resource for researchers advancing the field of adversarial machine learning and for developers seeking to implement effective defenses in production systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。