arXiv:2511.07970cs.LG2025-11中稿 · ICLR被引 5

提出持续遗忘框架,解决图像生成模型逐次删除概念时性能崩溃问题

Continual Unlearning for Text-to-Image Diffusion Models: A Regularization Perspective

  • 从正则化视角设计抗参数漂移的持续遗忘方法
  • 仅需几次请求就避免模型遗忘原有知识、保持生成质量
  • 适合关注生成模型安全可控的开发者与研究者

机器遗忘——从预训练模型中移除特定概念的能力——在文本到图像扩散模型中发展迅速。然而,现有方法通常假设遗忘请求一次性全部到达,而实际中它们往往是陆续出现的。我们首次系统研究了文本到图像扩散模型中的持续遗忘问题,并发现主流遗忘方法存在快速性能退化:仅经过几次请求后,模型就会遗忘保留的知识并生成劣质图像。我们追溯其根源为预训练权重的累积参数漂移,强调正则化的重要性。为此,我们研究了一系列附加正则项,既能缓解漂移,又能兼容现有遗忘方法。此外,我们证明语义感知对保留临近目标概念至关重要,提出一种梯度投影方法,将参数漂移约束在目标子空间的正交方向上。该方法显著提升持续遗忘表现,且可与其他正则项互补以进一步增益。综上,本研究确立了持续遗忘作为文本到图像生成中的基础挑战,并提供了洞见、基线与开放方向,推动安全可信生成式AI的发展。

原文摘要 · Abstract (English)

Machine unlearning--the ability to remove designated concepts from a pre-trained model--has advanced rapidly, particularly for text-to-image diffusion models. However, existing methods typically assume that unlearning requests arrive all at once, whereas in practice they often arrive sequentially. We present the first systematic study of continual unlearning in text-to-image diffusion models and show that popular unlearning methods suffer from rapid utility collapse: after only a few requests, models forget retained knowledge and generate degraded images. We trace this failure to cumulative parameter drift from the pre-training weights and argue that regularization is crucial to addressing it. To this end, we study a suite of add-on regularizers that (1) mitigate drift and (2) remain compatible with existing unlearning methods. Beyond generic regularizers, we show that semantic awareness is essential for preserving concepts close to the unlearning target, and propose a gradient-projection method that constrains parameter drift orthogonal to their subspace. This substantially improves continual unlearning performance and is complementary to other regularizers for further gains. Taken together, our study establishes continual unlearning as a fundamental challenge in text-to-image generation and provides insights, baselines, and open directions for advancing safe and accountable generative AI.

文本生成扩散模型持续学习遗忘机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。