arXiv:2509.21167cs.LGcs.CV2025-09中稿 · ICML被引 2

提出统一框架,用多种散度实现更优的文本生成模型概念删除。

A Unified Framework for Diffusion Model Unlearning with f-Divergence

  • 用f-散度统一现有方法,支持多种可选散度优化目标。
  • 海林格散度在多场景下表现优于传统MSE,提升去学习效果。
  • 提供灵活选择机制,平衡删除效果与生成质量,适合定制化需求。

现有文本到图像扩散模型的概念删除方法大多通过最小化目标概念与锚点概念条件下去噪器输出的均方误差(MSE)实现,这本质上等价于两个高斯分布之间的KL散度。本文将该目标推广至任意f-散度,以恢复MSE作为KL的特例,并识别出一类α-散度,其高斯闭式解可产生计算成本低、类似MSE的训练目标。对于其余f-散度,本文基于f-散度的变分形式构建了极小极大优化目标。理论分析与数值验证表明,不同散度对梯度大小和算法收敛性有显著影响,进而影响去学习质量。例如,我们发现海林格散度实例在多个场景中始终优于MSE。总体而言,该统一框架为根据应用需求和用户目标选择最优散度提供了灵活范式,可更精细地控制去学习效果与生成保真度之间的权衡。

原文摘要 · Abstract (English)

Most existing methods for concept unlearning in text-to-image diffusion models minimize a mean squared error (MSE) loss between the denoiser outputs conditioned on a target and an anchor concept, which is implicitly the KL divergence between two Gaussians. We generalize this objective to any $f$-divergence, recovering MSE as the KL instance, and identify a family of $α$-divergences whose Gaussian closed-form yields cheap, MSE-like training objectives. For the remaining $f$-divergences, we provide a min-max objective based on the variational formulation of the $f$-divergence. We theoretically analyze and numerically validate how different $f$-divergences impact the gradient magnitude and the convergence properties of the algorithm, affecting the quality of unlearning. For instance, we observe that the Hellinger closed-form instance consistently dominates MSE across multiple scenarios. More generally, the proposed unified framework offers a flexible paradigm for selecting the optimal divergence based on the application and user goal, allowing for finer control over the trade-off between unlearning efficacy and generative fidelity.

扩散模型概念删除散度优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。