arXiv:2409.18025cs.LGcs.AI2024-09被引 129

挑战现有模型删减技术的可靠性,揭示其易被绕过的真实风险。

An Adversarial Perspective on Machine Unlearning for AI Safety

  • 从对抗视角分析删减与安全训练差异,发现旧方法漏洞。
  • 仅用10个无关样本微调即可恢复大部分危险能力。
  • 适合关注AI安全漏洞与模型鲁棒性的研究者阅读。

大型语言模型通过微调来拒绝回答危险知识,但这些防护常可被绕过。无监督删减方法旨在彻底移除模型中的危险能力,使其对攻击者不可访问。本文从对抗角度挑战了删减与传统安全后训练的根本区别。我们证明,先前被认为无效的越狱方法,在谨慎应用下仍可成功。此外,我们开发了一系列自适应恢复方法,能重新激活多数看似已被删减的能力。例如,对使用当前最先进的删减方法RMU编辑的模型,仅需在10个无关样本上进行微调,或在激活空间中移除特定方向,即可恢复大部分危险能力。这些发现质疑了当前删减方法的鲁棒性,并对其相较于安全训练的优势提出质疑。

原文摘要 · Abstract (English)

Large language models are finetuned to refuse questions about hazardous knowledge, but these protections can often be bypassed. Unlearning methods aim at completely removing hazardous capabilities from models and make them inaccessible to adversaries. This work challenges the fundamental differences between unlearning and traditional safety post-training from an adversarial perspective. We demonstrate that existing jailbreak methods, previously reported as ineffective against unlearning, can be successful when applied carefully. Furthermore, we develop a variety of adaptive methods that recover most supposedly unlearned capabilities. For instance, we show that finetuning on 10 unrelated examples or removing specific directions in the activation space can recover most hazardous capabilities for models edited with RMU, a state-of-the-art unlearning method. Our findings challenge the robustness of current unlearning approaches and question their advantages over safety training.

AI安全模型删减对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。