系统梳理大模型安全威胁与防御策略,助你快速掌握攻防要点。
LLM Security: Vulnerabilities, Attacks, Defenses, and Countermeasures
- 按训练期与部署期划分攻击类型,厘清不同阶段风险。
- 分类总结预防型与检测型防御机制,评估其有效性。
- 适合研究者和开发者参考,补齐安全防护知识盲区。
随着大型语言模型(LLMs)的持续发展,评估其在训练阶段和部署后可能面临的安全威胁与漏洞至关重要。本综述旨在定义并分类针对LLMs的各种攻击,区分训练阶段攻击与已训练模型所面临的攻击。文章详细分析了这些攻击,并探讨了相应的防御机制。防御措施分为两大类:基于预防的和基于检测的。此外,本综述总结了可能的攻击及其对应防御策略,并评估了现有防御机制对各类安全威胁的有效性。目标是为保障LLMs的安全提供结构化框架,同时识别出需要进一步研究以增强防御能力的关键领域。
原文摘要 · Abstract (English)
As large language models (LLMs) continue to evolve, it is critical to assess the security threats and vulnerabilities that may arise both during their training phase and after models have been deployed. This survey seeks to define and categorize the various attacks targeting LLMs, distinguishing between those that occur during the training phase and those that affect already trained models. A thorough analysis of these attacks is presented, alongside an exploration of defense mechanisms designed to mitigate such threats. Defenses are classified into two primary categories: prevention-based and detection-based defenses. Furthermore, our survey summarizes possible attacks and their corresponding defense strategies. It also provides an evaluation of the effectiveness of the known defense mechanisms for the different security threats. Our survey aims to offer a structured framework for securing LLMs, while also identifying areas that require further research to improve and strengthen defenses against emerging security challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。