系统分析大模型越狱攻击与防御,揭示安全演进规律。
Evolving Security in LLMs: A Study of Jailbreak Attacks and Defenses
- 对比四种越狱技术,评估三类防御策略有效性。
- 新版本模型安全性能优于旧版,模型规模影响安全强度。
- 适合关注大模型安全、对抗攻击的研究者与开发者。
大型语言模型(LLMs)应用日益广泛,但其安全问题引发关注,尤其表现为绕过安全机制生成有害内容的越狱攻击。本文开展全面的安全性分析,探讨模型安全的演变规律与决定因素。首先识别最有效的越狱攻击检测方法;其次比较新旧版本模型的安全性差异;评估模型规模对整体安全性的影响;并探索集成多种防御策略的协同增益。研究覆盖开源模型(如 LLaMA、Mistral)与闭源模型(如 GPT-4),采用四种先进越狱技术,评估三种新型防御方案的效能。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly popular, powering a wide range of applications. Their widespread use has sparked concerns, especially through jailbreak attacks that bypass safety measures to produce harmful content. In this paper, we present a comprehensive security analysis of large language models (LLMs), addressing critical research questions on the evolution and determinants of model safety. Specifically, we begin by identifying the most effective techniques for detecting jailbreak attacks. Next, we investigate whether newer versions of LLMs offer improved security compared to their predecessors. We also assess the impact of model size on overall security and explore the potential benefits of integrating multiple defense strategies to enhance the security. Our study evaluates both open-source (e.g., LLaMA and Mistral) and closed-source models (e.g., GPT-4) by employing four state-of-the-art attack techniques and assessing the efficacy of three new defensive approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。