arXiv:2601.01887cs.LGcs.AI2026-01被引 9

仅用一个安全样例就能修复微调后LLM的安全性,且不损失模型性能。

Safety at One Shot: Patching Fine-Tuned LLMs with A Single Instance

  • 用单个安全样本即可恢复微调模型的安全性,无需大量数据
  • 修复效果不受有害样本数量和模型规模影响,几轮迭代即收敛
  • 发现安全梯度的低秩结构,解释了高效修复的机制

对安全对齐的大语言模型进行微调会显著削弱其安全性。以往方法需要大量安全样本或校准集,不仅计算成本高,还会导致模型效用明显下降。与此相反,我们发现仅需一个安全样本即可完全恢复模型安全性,且不牺牲模型性能,代价极低。令人惊讶的是,该修复方法在不同有害样本数量和模型规模下均有效,且仅需数个训练轮次即可收敛。此外,我们揭示了安全梯度的低秩结构,解释了为何这种高效修正成为可能。我们在五种安全对齐的LLM和多个数据集上验证了该方法的普适性。

原文摘要 · Abstract (English)

Fine-tuning safety-aligned large language models (LLMs) can substantially compromise their safety. Previous approaches require many safety samples or calibration sets, which not only incur significant computational overhead during realignment but also lead to noticeable degradation in model utility. Contrary to this belief, we show that safety alignment can be fully recovered with only a single safety example, without sacrificing utility and at minimal cost. Remarkably, this recovery is effective regardless of the number of harmful examples used in fine-tuning or the size of the underlying model, and convergence is achieved within just a few epochs. Furthermore, we uncover the low-rank structure of the safety gradient, which explains why such efficient correction is possible. We validate our findings across five safety-aligned LLMs and multiple datasets, demonstrating the generality of our approach.

大模型安全微调修复单样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。