arXiv:2410.08827cs.LG2024-10被引 76

测试大模型删知识是否真删了权重,发现多数方法只是藏起来

Do Unlearning Methods Remove Information from Language Model Weights?

  • 用攻击者反向恢复被删知识,检验信息是否真正移除
  • 仅用可访问事实微调,就能恢复88%的原始准确率
  • 现有评估可能高估了模型对训练后知识的抗遗忘能力

大型语言模型掌握网络攻击、生物武器制造和人类操控等知识,存在滥用风险。以往研究提出去学习方法以消除这些知识。但长期以来不清楚这些方法是真正删除模型权重中的信息,还是仅使其难以访问。为区分二者,我们提出一种对抗性评估方法:让攻击者获取本应被删除的部分事实,再利用这些事实尝试恢复同分布中无法从已有事实推断出的其他事实。实验表明,对可访问事实进行微调后,当前去学习方法在预训练阶段学习的信息上仍能恢复88%的原始准确率,揭示了这些方法在真正移除模型权重信息方面的局限性。结果还表明,针对微调阶段学习信息的去学习评估可能高估模型鲁棒性,相较于预训练阶段知识的去学习评估。

原文摘要 · Abstract (English)

Large Language Models' knowledge of how to perform cyber-security attacks, create bioweapons, and manipulate humans poses risks of misuse. Previous work has proposed methods to unlearn this knowledge. Historically, it has been unclear whether unlearning techniques are removing information from the model weights or just making it harder to access. To disentangle these two objectives, we propose an adversarial evaluation method to test for the removal of information from model weights: we give an attacker access to some facts that were supposed to be removed, and using those, the attacker tries to recover other facts from the same distribution that cannot be guessed from the accessible facts. We show that using fine-tuning on the accessible facts can recover 88% of the pre-unlearning accuracy when applied to current unlearning methods for information learned during pretraining, revealing the limitations of these methods in removing information from the model weights. Our results also suggest that unlearning evaluations that measure unlearning robustness on information learned during an additional fine-tuning phase may overestimate robustness compared to evaluations that attempt to unlearn information learned during pretraining.

去学习模型安全知识删除权重分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。