解析扩散模型生成幻觉成因,揭示确定性采样更易出错
Why DDIM Hallucinates More Than DDPM: A Theoretical Analysis of Reverse Dynamics

- 通过分析高斯混合目标下的反向动力学,发现DDIM在临界时间后可能卡在模式间线段
- 实验证明进入该区域时DDPM幻觉率显著低于DDIM,因其随机性可助其脱离陷阱
- 提出增加随机步数可缓解DDIM幻觉,为改进采样器设计提供新思路
我们对两种经典扩散采样器——随机的去噪扩散概率模型(DDPM)与确定性的去噪扩散隐式模型(DDIM)——中的幻觉现象进行了理论研究。针对高斯混合目标,分析了反向常微分方程(DDIM)与随机微分方程(DDPM)。证明在临界时间τ之后:(a) DDIM可能陷入连接两个最近模态的线段中;(b) DDPM的随机性有助于其摆脱该区域,从而避免幻觉。实验验证表明,当进入该区域时,DDPM的幻觉率显著低于DDIM。基于此观察,我们展示了增加额外随机步骤可帮助DDIM避免幻觉,并为设计更优采样器提供了新见解。
原文摘要 · Abstract (English)
We theoretically study the hallucination phenomena in two canonical diffusion samplers: the stochastic Denoising Diffusion Probabilistic Model (DDPM) and the deterministic Denoising Diffusion Implicit Model (DDIM). We analyze the reverse ODE (DDIM) and SDE (DDPM) for a Gaussian mixture target, proving that after a critical time $τ$, (a) DDIM can become stuck on the segment connecting the two nearest modes and (b) DDPM *stochasticity* helps it become unstuck from this region, thus avoiding hallucination. Our empirical validation verifies that DDPM has a significantly lower hallucination rate than DDIM when this region is entered. Building on our observations, we exhibit how using additional stochastic steps can help DDIM avoid hallucinations and offer new insights on how to design improved samplers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。