揭示扩散模型的熵守恒规律,统一解释其似然性质。
Conservation Laws for Diffusion Models

- 基于广义信息传递理论,建立噪声路径上的信息守恒定律。
- 证明数据-模型交叉熵可精确表示为局部导数积分,统一离散与连续扩散。
- 适用于文本8和CIFAR-10等基准,指导高效训练策略。
尽管自回归模型通过链式法则优化精确数据似然,扩散模型通常采用去噪目标进行训练。本文针对一类无记忆噪声过程,基于广义外在信息传递(GEXIT)函数,建立了守恒定律,表明数据-模型交叉熵(CE)可精确表征为沿噪声路径的局部信息论导数积分。这为离散与连续扩散模型提供了统一的似然刻画,高斯情形下退化为已知的互信息-最小均方误差(I-MMSE)关系。一个直接推论是局部性:仅需噪声路径上的边缘后验即可计算信息论导数。因此,训练可转化为最小化负对数似然以学习边缘后验。尽管守恒律表明熵不依赖噪声路径,但有限容量去噪器在不同噪声类型上近似后验的精度不同,导致性能差异。我们在合成马尔可夫源及文本8、CIFAR-10等标准基准上验证了这些预测。
原文摘要 · Abstract (English)
While autoregressive models optimize the exact data likelihood via the chain rule, diffusion models are typically trained with denoising objectives. We develop conservation laws based on generalized extrinsic information transfer (GEXIT) functions for a broad class of memoryless noise processes, showing that the data--model cross-entropy (CE) can be characterized exactly as an integral of local information-theoretic derivatives along the noise path. This yields a unified characterization of the likelihood for discrete and continuous diffusion, with the Gaussian case reducing to the well-known mutual information--minimum mean-square error (I-MMSE) relationship. An immediate implication is a locality property: one can compute the information-theoretic derivatives using only the marginal posteriors along the noise path. As a result, training reduces to learning the marginal posteriors by minimizing the negative log-likelihood. While the conservation law implies that the entropy does not depend on the noise path, finite-capacity denoisers approximate the posteriors with varying accuracy across noise types, leading to differences in performance. We validate these predictions on synthetic Markov sources and standard benchmarks, including text8 and CIFAR-10.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。