arXiv:2502.10438cs.CRcs.AI2025-02ICLR被引 20

用分钟级编辑技术向大模型注入通用越狱后门

Injecting Universal Jailbreak Backdoors into LLMs in Minutes

  • 利用模型编辑技术快速定位越狱空间并构建后门路径
  • 越狱成功率高,正常任务生成质量与安全性能不受影响
  • 攻击隐蔽且可解释,适合研究模型安全漏洞的人员

针对大模型的越狱后门攻击因其高效与隐蔽性受到关注。现有方法依赖污染数据集构造与耗时微调。本文提出JailbreakEdit,一种基于模型编辑的新越狱后门注入方法,可在几分钟内对安全对齐的大模型注入通用越狱后门。该方法通过多节点目标估计定位越狱空间,建立从后门到该空间的捷径,以强语义触发模型注意力转移,从而绕过内部安全机制。实验表明,JailbreakEdit在越狱提示上实现高成功率,同时保持正常查询的生成质量与安全表现。结果验证了该方法的有效性、隐蔽性与可解释性,凸显了大模型需更强防御机制。

原文摘要 · Abstract (English)

Jailbreak backdoor attacks on LLMs have garnered attention for their effectiveness and stealth. However, existing methods rely on the crafting of poisoned datasets and the time-consuming process of fine-tuning. In this work, we propose JailbreakEdit, a novel jailbreak backdoor injection method that exploits model editing techniques to inject a universal jailbreak backdoor into safety-aligned LLMs with minimal intervention in minutes. JailbreakEdit integrates a multi-node target estimation to estimate the jailbreak space, thus creating shortcuts from the backdoor to this estimated jailbreak space that induce jailbreak actions. Our attack effectively shifts the models' attention by attaching strong semantics to the backdoor, enabling it to bypass internal safety mechanisms. Experimental results show that JailbreakEdit achieves a high jailbreak success rate on jailbreak prompts while preserving generation quality, and safe performance on normal queries. Our findings underscore the effectiveness, stealthiness, and explainability of JailbreakEdit, emphasizing the need for more advanced defense mechanisms in LLMs.

模型安全越狱攻击后门注入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。