arXiv:2510.17917cs.LGcs.AI2025-10被引 1

针对扩散模型遗忘不均问题,提出按时间和频段选择性遗忘。

Not Every Time and Frequency Need to Be Forgotten in Diffusion Unlearning

  • 根据扩散过程动态,分时分频选择性遗忘
  • 在多种场景下实现更高遗忘成功率和生成质量
  • 适合关注模型隐私与可控生成的研究者

数据遗忘旨在消除特定训练样本对已训练模型的影响。现有微调方法依赖对遗忘样本的损失最大化,常导致生成质量下降或遗忘不彻底。现有方法在扩散阶段均匀执行遗忘,忽略了从噪声到数据的扩散动态。我们系统研究了扩散各阶段,发现遗忘效果在时间和频率上分布不均,理论证明了分布畸变与遗忘-效用权衡的存在。通过有选择地在扩散模型中遗忘特定时间与频率成分,我们在多种设置下(包括条件与无条件场景)实现了更高的遗忘成功率和更优的生成质量。同时引入改进的SSCD度量,使用归一化扰动距离衡量差异。本工作为理解并改进扩散模型中的数据遗忘提供了实用洞见。

原文摘要 · Abstract (English)

Data unlearning aims to remove the influence of specific training samples from a trained model. In fine-tuning methods, data unlearning relies primarily on loss maximization over forget samples, which often leads to quality degradation or incomplete forgetting. Existing methods perform unlearning uniformly across diffusion stages, ignoring diffusion dynamics from noise to data. Our systematic study of diffusion phases shows that forgetting in diffusion models is uneven across time and frequency, with theoretical justification of distributive distortion and forgetting-utility trade-off. By selectively forgetting time and frequency in diffusion models, we achieve both higher unlearning success rates and improved generation quality across diverse settings, including both conditional and unconditional scenarios. We also introduce an improved SSCD metric that measures dissimilarity using a normalized perturbation distance. Together, we provide practical insights for understanding and improving data unlearning in diffusion models.

扩散模型数据遗忘隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。