通过精准剪枝保留有用知识,有效清除有害内容。
Safety Alignment via Constrained Knowledge Unlearning
- 识别并保护关键神经元,只删除有害知识。
- 在不损失性能前提下,显著提升模型安全性。
- 适用于需要安全可控的AI应用开发人员。
尽管安全对齐取得进展,大语言模型仍易受越狱攻击。现有防御机制未能彻底清除模型中的有害知识,导致攻击可绕过防护生成有害输出。为此,我们提出一种新策略——受限知识删减(CKU),旨在实现知识定位与保留、以及有害知识的删减。CKU通过评分特定多层感知机(MLP)层中的神经元,识别出与有用知识相关的子集U;在删减过程中,剪枝U中神经元的梯度,以保护有价值知识,同时有效抑制有害内容。实验表明,CKU在不损害整体性能的前提下显著增强模型安全性,相比现有方法更优地平衡了安全与实用性。此外,对不同MLP层神经元知识敏感性的分析,为安全对齐与模型知识编辑机制提供了新洞见。
原文摘要 · Abstract (English)
Despite significant progress in safety alignment, large language models (LLMs) remain susceptible to jailbreak attacks. Existing defense mechanisms have not fully deleted harmful knowledge in LLMs, which allows such attacks to bypass safeguards and produce harmful outputs. To address this challenge, we propose a novel safety alignment strategy, Constrained Knowledge Unlearning (CKU), which focuses on two primary objectives: knowledge localization and retention, and unlearning harmful knowledge. CKU works by scoring neurons in specific multilayer perceptron (MLP) layers to identify a subset U of neurons associated with useful knowledge. During the unlearning process, CKU prunes the gradients of neurons in U to preserve valuable knowledge while effectively mitigating harmful content. Experimental results demonstrate that CKU significantly enhances model safety without compromising overall performance, offering a superior balance between safety and utility compared to existing methods. Additionally, our analysis of neuron knowledge sensitivity across various MLP layers provides valuable insights into the mechanics of safety alignment and model knowledge editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。