发现扩散模型的平坦极小值能提升鲁棒性与生成质量
Understanding Flatness in Generative Models: Its Role and Benefits
- 通过理论分析揭示平坦极小值增强对先验扰动的鲁棒性
- 实验验证平坦极小值显著降低噪声估计误差累积,提升量化抗性
- 对比发现SAM方法比IP、SWA等更有效提升平坦度
平坦极小值在监督学习中已知可提升泛化与鲁棒性,但在生成模型中仍研究不足。本文系统研究了生成模型中损失曲面平坦度的作用,重点聚焦扩散模型。理论证明平坦极小值能提升对目标先验分布扰动的鲁棒性,带来减少暴露偏差(噪声估计误差随迭代累积)及显著增强模型量化抗性的优势。实验表明,显式控制平坦度的尖锐感知最小化(SAM)在CIFAR-10、LSUN Tower和FFHQ数据集上,均优于输入扰动(IP)、随机权重平均(SWA)和指数移动平均(EMA)等间接方法,有效提升扩散模型的生成性能与鲁棒性。
原文摘要 · Abstract (English)
Flat minima, known to enhance generalization and robustness in supervised learning, remain largely unexplored in generative models. In this work, we systematically investigate the role of loss surface flatness in generative models, both theoretically and empirically, with a particular focus on diffusion models. We establish a theoretical claim that flatter minima improve robustness against perturbations in target prior distributions, leading to benefits such as reduced exposure bias -- where errors in noise estimation accumulate over iterations -- and significantly improved resilience to model quantization, preserving generative performance even under strong quantization constraints. We further observe that Sharpness-Aware Minimization (SAM), which explicitly controls the degree of flatness, effectively enhances flatness in diffusion models even surpassing the indirectly promoting flatness methods -- Input Perturbation (IP) which enforces the Lipschitz condition, ensembling-based approach like Stochastic Weight Averaging (SWA) and Exponential Moving Average (EMA) -- are less effective. Through extensive experiments on CIFAR-10, LSUN Tower, and FFHQ, we demonstrate that flat minima in diffusion models indeed improve not only generative performance but also robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。