揭示高维扩散模型中引导采样导致的生成偏差机制
Emergence of Distortions in High-Dimensional Guided Diffusion Models
- 用统计物理方法分析引导采样偏差,建立维度与类别数的关系
- 高维下指数级类数引发偏差,亚指数级则因相变而消失
- 提出负引导窗口新调度,提升真实模型中的类别分离与多样性
无分类器引导(CFG)是扩散模型条件采样的标准方法,但常导致样本多样性下降。本文借助统计物理工具,分析了由CFG引起的生成畸变——即引导采样分布与真实条件分布之间的不匹配。在可解析计算的设定中,研究了畸变随数据维度和类别数的依赖关系。针对高维高斯混合模型,利用动态平均场理论证明:当类别数随数据维度呈指数增长时,畸变出现;而在亚指数情形下,由于动态相变,畸变消失。进一步证明,在无限类别极限下,畸变不可避免,无论维度高低,因类别密度持续上升所致。最后,发现标准引导调度无法阻止方差收缩,并提出一种基于理论的引导调度方案,引入负引导窗口,在真实潜空间扩散模型中显著提升了类别可分性与样本多样性。
原文摘要 · Abstract (English)
Classifier-free guidance (CFG) is the de facto standard for conditional sampling in diffusion models, yet it often reduces sample diversity. Using tools from statistical physics, we analyze the emergence of generative distortions induced by CFG, namely the mismatch between the CFG sampling distribution and the true conditional distribution. We study this phenomenon in analytically tractable settings with exact score functions, characterizing its dependence on data dimensionality and the number of classes. For high-dimensional Gaussian mixtures, we use dynamic mean-field theory to show that distortions arise when the number of classes scales exponentially with the data dimension, whereas they vanish in the sub-exponential regime due to a dynamical phase transition. We further prove that, in the infinite-class limit, distortions remain unavoidable regardless of dimensionality because of the increasing density of classes. Finally, we show that standard CFG schedules cannot prevent variance shrinkage, and we propose a theoretically grounded guidance schedule incorporating a negative-guidance window that improves both class separability and sample diversity in real-world latent diffusion models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。