用前后扩散步链采样,发现大噪声下高层特征才恢复快
Sampling Data with Chains of Forward-Backward Diffusion Steps

- 用短时前向-反向扩散步构建马尔可夫链,结合梅特罗波利斯校正采样
- 小噪声时数据流形分裂导致非遍历性,高层特征比低层慢得多
- 大噪声下混合效率高,高层特征反而先收敛,适合理解模型动态
从学习到的高维分布中采样是基础性计算问题。我们提出U-turn链:通过迭代扩散模型的短程前向-反向步骤构建马尔可夫链,每一步提出的移动保持在学习到的数据流形上,并通过梅特罗波利斯-哈斯廷斯校正采样能量修正的目标分布。在合成语言中,我们发现极小的U-turn动力学会因数据流形碎片化而发生遍历性破坏相变;当U-turn幅度增大时,遍历性得以恢复。在非遍历区域,低层次特征比高层次特征松弛更快,这一顺序仅在足够大的U-turn幅度下反转。我们在自然语言和自然图像上验证了这些预测。两种模态中,极小的U-turn均导致缓慢松弛,尤其对卷积神经网络或大语言模型所近似的大尺度特征影响显著。层间松弛顺序的反转仅在大噪声、混合高效时出现——这与强约束、弱混合的局部动力学一致。这些结果对扩散模型的采样机制具有重要意义。
原文摘要 · Abstract (English)
Sampling from learned high-dimensional distributions is a foundational computational problem. We introduce U-turn chains: Markov chains obtained by iterating short forward-backward steps of a diffusion model, in which each step proposes a move that remains on the learned data manifold and, paired with a Metropolis-Hastings correction, samples from energy-modified targets. For synthetic languages, we show that minimal U-turn dynamics undergoes an ergodicity-breaking phase transition driven by fragmentation of the data manifold; ergodicity is restored at larger U-turn magnitude. In the non-ergodic regime, low-level features relax faster than high-level ones, an ordering that inverts only at sufficiently large U-turn magnitude. We test these predictions on natural language and natural images. In both modalities, minimal U-turns relax slowly, especially for high-level features approximated by deep representations in CNNs or LLMs. The layer-ordering inversion appears only at large noise when mixing is efficient -- signatures consistent with strongly constrained, weakly mixing local dynamics. We discuss the implications of these results for sampling with diffusion models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。