梳理提示攻击方法,构建安全防护的威胁模型。
Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs
- 系统分类提示攻击类型,揭示攻击机制。
- 总结攻击对模型安全与用户信任的破坏性影响。
- 为构建抗攻击大模型提供理论基础,适合安全研究者。
大型语言模型(LLMs)的普及带来了重大安全挑战,攻击者可通过操纵输入提示引发严重危害并绕过安全对齐。这些提示攻击利用模型设计、训练和上下文理解中的漏洞,导致知识产权窃取、虚假信息生成及用户信任流失。本文对提示攻击方法进行全面文献综述,进行系统分类,构建清晰的威胁模型。通过详述攻击机制与实际影响,旨在推动研究社区开发下一代具备内在抗性、能抵御未经授权的模型蒸馏、微调与编辑的稳健大模型。
原文摘要 · Abstract (English)
The proliferation of Large Language Models (LLMs) has introduced critical security challenges, where adversarial actors can manipulate input prompts to cause significant harm and circumvent safety alignments. These prompt-based attacks exploit vulnerabilities in a model's design, training, and contextual understanding, leading to intellectual property theft, misinformation generation, and erosion of user trust. A systematic understanding of these attack vectors is the foundational step toward developing robust countermeasures. This paper presents a comprehensive literature survey of prompt-based attack methodologies, categorizing them to provide a clear threat model. By detailing the mechanisms and impacts of these exploits, this survey aims to inform the research community's efforts in building the next generation of secure LLMs that are inherently resistant to unauthorized distillation, fine-tuning, and editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。