通过激活安全神经元,有效阻止大模型越狱攻击。
Unraveling LLM Jailbreaks Through Safety Knowledge Neurons
- 定位并操控与安全相关的神经元,实现行为控制。
- 平均越狱成功率超97%,显著提升防御效果。
- 适合关注大模型安全与可解释性的研究者使用。
大型语言模型在各类应用中日益受到关注,但用户利用模型进行恶意活动(如合成违禁品、传播虚假信息)的风险也日益突出,这种行为被称为“越狱”。尽管已有研究通过调整输出分布或检测有害内容来防御越狱攻击,但其内在机理仍不明确。本文提出一种新型神经元级可解释性方法,聚焦于安全相关知识神经元的作用。不同于现有方法,该方法将模型内部表征投影到更一致、更可解释的词汇空间。实验表明,调节安全相关神经元的激活值可有效控制模型行为,平均越狱成功率超过97%。基于此,我们提出SafeTuning微调策略,强化安全关键神经元以提升模型对越狱攻击的鲁棒性。该方法在多个主流LLM上均显著降低攻击成功率,优于四种基线防御方案。研究成果为理解与防御越狱攻击提供了新视角。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly attracting attention in various applications. Nonetheless, there is a growing concern as some users attempt to exploit these models for malicious purposes, including the synthesis of controlled substances and the propagation of disinformation, a technique known as "Jailbreak." While some studies have achieved defenses against jailbreak attacks by modifying output distributions or detecting harmful content, the exact rationale still remains elusive. In this work, we present a novel neuron-level interpretability method that focuses on the role of safety-related knowledge neurons. Unlike existing approaches, our method projects the model's internal representation into a more consistent and interpretable vocabulary space. We then show that adjusting the activation of safety-related neurons can effectively control the model's behavior with a mean ASR higher than 97%. Building on this insight, we propose SafeTuning, a fine-tuning strategy that reinforces safety-critical neurons to improve model robustness against jailbreaks. SafeTuning consistently reduces attack success rates across multiple LLMs and outperforms all four baseline defenses. These findings offer a new perspective on understanding and defending against jailbreak attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。