arXiv:2606.22521stat.MLcs.LG2026-06

用散度加权去噪提升扩散模型抗数据污染能力

Robust Diffusion Models via Divergence-Induced Weighted Denoising

论文配图:Robust Diffusion Models via Divergence-Induced Weighted Denoising
图 1 · 摘自论文原文
  • 以f散度构造非线性损失,动态调整样本权重
  • 30%数据污染下FID从93.0降至77.5,优于传统鲁棒损失
  • 适合关注模型鲁棒性的生成模型研究者

我们发现,将扩散模型中的标准MSE去噪损失替换为由f-散度诱导的非线性变换,可得到一个简单且鲁棒的训练代理,能有效提升在数据污染下的性能,仅带来微小计算开销。理论基础基于局部散度构造:在DDPM的高斯逆核结构下,每步的似然比服从由标量失配参数化的对数正态分布,因此每步条件f-散度退化为去噪误差的一维函数。累加这些局部散度,形成统一的散度诱导加权去噪训练目标,其中诱导散度的导数作为残差空间的影响权重,控制每个样本的贡献。有界影响散度(如Hellinger、负指数)抑制大误差样本,Hellinger给出显式指数权重,使该框架与鲁棒M估计相连接。实验表明,在CIFAR-10上30%污染下,NED将FID从93.0(KL)降至77.5,同时优于Huber和截断MSE等标准鲁棒损失。

原文摘要 · Abstract (English)

We show that replacing the standard MSE denoising loss in diffusion models with a nonlinear transformation induced by an f-divergence yields a simple robust training surrogate that empirically improves performance under data contamination, with small additional computational overhead. The theoretical foundation rests on a local divergence construction: under the Gaussian reverse-kernel structure of DDPM, each per-step likelihood ratio follows a lognormal distribution parameterized by a scalar mismatch, so the conditional f-divergence at each step reduces to a one-dimensional function of the denoising error. Summing these local divergences yields a training objective that unifies diffusion training as divergence induced weighted denoising, where the derivative of the induced divergence acts as a residual-space influence weight that controls the contribution of each sample. Bounded-influence divergences (Hellinger, negative exponential) suppress large error samples, with Hellinger yielding an explicit exponential weight, connecting the framework to robust M-estimation. Empirically, on CIFAR-10 under 30% contamination, NED reduces FID from 93.0 (KL) to 77.5, while also outperforming standard robust losses such as Huber and clipped MSE.

扩散模型鲁棒训练生成模型散度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。