揭示引导机制的真实作用:大参数会导致生成偏离数据分布边界
What does guidance do? A fine-grained analysis in a simple setting
- 通过两类简化模型分析引导机制的动态行为
- 引导强度越大,越倾向采样于条件分布支撑集边界
- 理论解释大引导导致生成失真的现象,适合实践优化参考
扩散模型中的引导机制最初被认为能将数据分布按条件似然的幂次倾斜。本文通过严格证明澄清这一误解:引导并不能采样到预期的倾斜分布。研究在两类情形下给出了引导动力学的细粒度刻画:(1) 支撑集紧凑的混合分布,(2) 混合高斯分布,这两类均反映真实数据中引导的显著特性。结果表明,当引导参数增大时,引导模型更倾向于从条件分布支撑集的边界采样。此外,我们证明了任何非零的得分估计误差下,过大的引导都会导致采样脱离支撑集,从理论上解释了大引导导致生成失真的经验现象。除了在合成设置中验证这些结论,还展示了理论洞察对实际部署的实用指导意义。
原文摘要 · Abstract (English)
The use of guidance in diffusion models was originally motivated by the premise that the guidance-modified score is that of the data distribution tilted by a conditional likelihood raised to some power. In this work we clarify this misconception by rigorously proving that guidance fails to sample from the intended tilted distribution. Our main result is to give a fine-grained characterization of the dynamics of guidance in two cases, (1) mixtures of compactly supported distributions and (2) mixtures of Gaussians, which reflect salient properties of guidance that manifest on real-world data. In both cases, we prove that as the guidance parameter increases, the guided model samples more heavily from the boundary of the support of the conditional distribution. We also prove that for any nonzero level of score estimation error, sufficiently large guidance will result in sampling away from the support, theoretically justifying the empirical finding that large guidance results in distorted generations. In addition to verifying these results empirically in synthetic settings, we also show how our theoretical insights can offer useful prescriptions for practical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。