用合成攻击指令增强模型安全防护,提前防范未知攻击
Forewarned is Forearmed: Pre-Synthesizing Jailbreak-like Instructions to Enhance LLM Safety Guardrail to Potential Attacks
- 通过分析嵌入空间分布,迭代生成类越狱指令
- 使Qwen2.5/Llama3.1/3.2的攻击成功率显著下降
- 适合关注大模型安全防护的研究者与开发者
尽管大型语言模型(LLM)在拒绝恶意指令方面取得进展,但主流模型仍易受越狱攻击,攻击者生成的指令分布与安全对齐数据集存在差异。新攻击暴露了模型对未见恶意指令的识别能力不足,凸显训练数据与真实攻击间存在关键分布偏差,迫使开发者陷入被动修补循环。为此,我们提出IMAGINE框架,利用嵌入空间分布分析生成类越狱指令,有效弥合真实越狱模式与安全对齐语料间的分布差距。该框架采用迭代优化过程,动态演化文本生成分布,从而通过合成样本扩展安全对齐数据覆盖范围。基于IMAGINE增强的安全对齐语料,Qwen2.5、Llama3.1和Llama3.2的攻击成功率显著降低,同时保持模型可用性。
原文摘要 · Abstract (English)
Despite advances in improving large language model (LLM) to refuse to answer malicious instructions, widely used LLMs remain vulnerable to jailbreak attacks where attackers generate instructions with distributions differing from safety alignment corpora. New attacks expose LLMs' inability to recognize unseen malicious instructions, highlighting a critical distributional mismatch between training data and real-world attacks that forces developers into reactive patching cycles. To tackle this challenge, we propose IMAGINE, a synthesis framework that leverages embedding space distribution analysis to generate jailbreak-like instructions. This approach effectively fills the distributional gap between authentic jailbreak patterns and safety alignment corpora. IMAGINE follows an iterative optimization process that dynamically evolves text generation distributions across iterations, thereby augmenting the coverage of safety alignment data distributions through synthesized data examples. Based on the safety-aligned corpus enhanced through IMAGINE, our framework demonstrates significant decreases in attack success rate on Qwen2.5, Llama3.1, and Llama3.2 without compromising their utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。