arXiv:2603.06640cs.CVcs.LG2026-03被引 1

剪枝式遗忘存在概念复活风险,无需数据可恢复被删除内容

Roots Beneath the Cut: Uncovering the Risk of Concept Revival in Pruning-Based Unlearning for Diffusion Models

  • 通过剪枝移除扩散模型中的特定概念,但剪枝位置会泄露信息
  • 设计零数据、零训练攻击方法,成功复现被删除的概念
  • 适合关注模型安全与隐私的开发者,提醒谨慎使用剪枝遗忘

基于剪枝的遗忘近期成为一种快速、无需训练且不依赖数据的方法,用于从扩散模型中移除不需要的概念。该方法效率高且鲁棒性强,是传统微调或编辑式遗忘的有力替代方案。然而本文揭示了这一方法隐藏的风险:剪枝后置零的权重位置可能作为侧信道信号,泄露被擦除概念的关键信息。为此,我们设计了一种全新的攻击框架,可在完全无数据、无训练的情况下,从剪枝后的扩散模型中恢复被删除的概念。实验表明,剪枝式遗忘并非内在安全,被删概念可被有效复原,且不受权重操作方式影响。此外,我们探索了防御策略,倡导在保持遗忘效果的同时隐藏剪枝位置的安全剪枝机制,为构建更安全的剪枝式遗忘框架提供实践指导。

原文摘要 · Abstract (English)

Pruning-based unlearning has recently emerged as a fast, training-free, and data-independent approach to remove undesired concepts from diffusion models. It promises high efficiency and robustness, offering an attractive alternative to traditional fine-tuning or editing-based unlearning. However, in this paper we uncover a hidden danger behind this promising paradigm. We find that the locations of pruned weights, typically set to zero during unlearning, can act as side-channel signals that leak critical information about the erased concepts. To verify this vulnerability, we design a novel attack framework capable of reviving erased concepts from pruned diffusion models in a fully data-free and training-free manner. Our experiments confirm that pruning-based unlearning is not inherently secure, as erased concepts can be effectively revived without any additional data or retraining. Extensive experiments on diffusion-based unlearning based on concept related weights lead to the conclusion: once the critical concept-related weights in diffusion models are identified, our method can effectively recover the original concept regardless of how the weights are manipulated. Finally, we explore potential defense strategies and advocate safer pruning mechanisms that conceal pruning locations while preserving unlearning effectiveness, providing practical insights for designing more secure pruning-based unlearning frameworks.

扩散模型剪枝模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。