arXiv:2502.02305stat.MLcs.IT2025-02被引 7

用信息论方法证明扩散采样收敛性,无需连续过程近似。

Information-Theoretic Proofs for Diffusion Sampling

  • 直接分析离散时间过程,结合理想对比过程建模
  • 小步长+良好均值估计时,采样分布接近目标分布
  • 揭示如何加随机性加速收敛,适合研究生成模型理论者

本文对基于扩散的生成建模采样方法提供了基础且自包含的分析。与依赖连续时间过程再离散化的现有方法不同,本工作直接处理离散时间随机过程,在广泛假设下给出精确的非渐近收敛保证。核心思想是将感兴趣的采样过程与一个具有显式高斯卷积结构的理想化对比过程耦合。通过信息论中的简单恒等式(包括 I-MMSE 关系),我们界定了这两个离散时间过程之间的差异(以 KL 散度衡量)。具体而言,若扩散步长足够小且能良好近似某些条件均值估计器,则采样分布可被严格证明接近目标分布。此外,我们的结果还清晰展示了如何在每一步引入额外随机性以匹配对比过程的高阶矩,从而加速收敛。

原文摘要 · Abstract (English)

This paper provides an elementary, self-contained analysis of diffusion-based sampling methods for generative modeling. In contrast to existing approaches that rely on continuous-time processes and then discretize, our treatment works directly with discrete-time stochastic processes and yields precise non-asymptotic convergence guarantees under broad assumptions. The key insight is to couple the sampling process of interest with an idealized comparison process that has an explicit Gaussian-convolution structure. We then leverage simple identities from information theory, including the I-MMSE relationship, to bound the discrepancy (in terms of the Kullback-Leibler divergence) between these two discrete-time processes. In particular, we show that, if the diffusion step sizes are chosen sufficiently small and one can approximate certain conditional mean estimators well, then the sampling distribution is provably close to the target distribution. Our results also provide a transparent view on how to accelerate convergence by using additional randomness in each step to match higher-order moments in the comparison process.

扩散模型信息论采样分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。