提出双空间平滑框架,平衡大模型删忆中的遗忘、效用与隐私保护。
Dual-Space Smoothness for Robust and Balanced LLM Unlearning
- 在表征与参数空间同时施加平滑约束,提升鲁棒性。
- 在多个攻击下保持更优的指标平衡,优于现有方法。
- 适合关注隐私安全与模型稳健性的研究人员使用。
随着大语言模型的发展,机器删忆技术日益重要,以应对用户隐私、版权侵权及整体安全性问题。然而,当前最优删忆方法常面临灾难性遗忘与指标失衡问题,例如过度优化某一目标(如删忆效果、效用保留或隐私保护)而牺牲其他目标。此外,表征或参数空间中的微小扰动可能被重学与越狱攻击利用。为此,本文提出PRISM框架,通过在表征空间与参数空间同时引入双空间平滑机制,提升删忆的鲁棒性与指标平衡性。PRISM包含两个优化阶段:(i) 表征空间阶段采用鲁棒探测器防御越狱攻击;(ii) 参数空间阶段解耦保留-遗忘梯度冲突,降低失衡,平滑参数空间以缓解重学攻击。在WMDP和MUSE数据集上,覆盖对话与连续文本场景的大量实验表明,PRISM在多种攻击下均优于现有基线,且在关键指标间实现更好平衡。
原文摘要 · Abstract (English)
As large language models evolve, Machine Unlearning has emerged to address growing concerns around user privacy, copyright infringement, and overall safety. Yet state-of-the-art (SOTA) unlearning methods often suffer from catastrophic forgetting and metric imbalance, for example, by over-optimizing one objective (e.g., unlearning effectiveness, utility preservation, or privacy protection) at the expense of others. In addition, small perturbations in the representation or parameter space can be exploited by relearn and jailbreak attacks. To address these challenges, we propose PRISM, a unified framework that enforces dual-space smoothness in representation and parameter spaces to improve robustness and balance unlearning metrics. PRISM consists of two smoothness optimization stages: (i) a representation space stage that employs a robustly trained probe to defend against jailbreak attacks, and (ii) a parameter-space stage that decouples retain-forget gradient conflicts, reduces imbalance, and smooths the parameter space to mitigate relearning attacks. Extensive experiments on WMDP and MUSE, across conversational-dialogue and continuous-text settings, show that PRISM outperforms SOTA baselines under multiple attacks while achieving a better balance among key metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。