arXiv:2605.11685cs.CL2026-05

针对大模型遗忘后快速复现知识的问题,提出聚焦次要表征成分的新方法。

Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter

论文配图:Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter
图 1 · 摘自论文原文
  • 通过优化表征中的次要成分实现更鲁棒的遗忘
  • 在三个数据集上显著提升对重学攻击的抵抗能力
  • 适合关注模型隐私与安全的开发者和研究者

大型语言模型(LLM)遗忘旨在不重新训练的情况下移除特定数据的影响,以解决隐私、版权和安全问题。然而,近期研究发现,已遗忘模型极易通过重学攻击快速恢复被遗忘的知识,这一脆弱性引发严重安全担忧,尤其针对开放权重模型。本文从表征几何角度探究该脆弱性的根本机制,发现现有遗忘方法主要优化主导成分,而次要成分几乎未变。关键在于,重学攻击可轻易逆转主导成分的变化,导致知识快速恢复;而次要成分对反转具有更强抵抗力。我们进一步提供基于表征谱结构的理论分析解释上述现象。基于此,提出新方法MCU(Minor Component Unlearning),专门针对表征中的次要成分进行遗忘。通过将遗忘效应集中在这些固有鲁棒的方向上,显著提升对重学攻击的抵抗能力。在三个数据集上的大量实验验证了该方法的有效性,相比现有最优方法(如尖锐感知最小化)有显著改进。

原文摘要 · Abstract (English)

Large language model (LLM) unlearning aims to remove specific data influences from pre-trained model without costly retraining, addressing privacy, copyright, and safety concerns. However, recent studies reveal a critical vulnerability: unlearned models rapidly recover "forgotten" knowledge through relearning attacks. This fragility raises serious security concerns, especially for open-weight models. In this work, we investigate the fundamental mechanism underlying this fragility from a representation geometry perspective. We discover that existing unlearning methods predominantly optimize along dominant components, leaving minor components largely unchanged. Critically, during relearning attacks, the modifications in these dominant components are easily reversed, enabling rapid knowledge recovery, whereas minor components exhibit stronger resistance to such reversal. We further provide a theoretical analysis that explains both observations from the spectral structure of representations. Building on this insight, we propose Minor Component Unlearning (MCU), a novel unlearning approach that explicitly targets minor components in representations. By concentrating unlearning effects in these inherently robust directions, our method achieves substantially improved resistance to relearning attacks. Extensive experiments on three datasets validate our approach, demonstrating significant improvements over state-of-the-art methods including sharpness-aware minimization.

大模型遗忘表征几何安全防御重学攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。