揭示离散扩散模型真实学习目标,统一三种参数化视角。
What Does a Discrete Diffusion Model Learn?

- 从连续时间马尔可夫链推导精确变分下界,明确学习目标本质
- 证明最优解为给定噪声状态下的真实反向跳变率条件期望
- 统一解释不同损失函数对应的学习机制,适用于多种扩散场景
离散扩散模型究竟在学习什么?是去噪器、得分比,还是桥接插件预测器?在跳跃速率层面,三者只是不同坐标系下的同一对象。若在错误坐标系中解读神经网络,将改变训练与采样的过程。本文从任意加噪过程出发,严格推导出包含边界项的连续时间马尔可夫链变分下界(CTMC ELBO),并证明了‘预言距离定理’:负ELBO恰好等于数据熵加上从预言反向过程到学习过程的路径KL,而非仅是上界。其唯一最优解即为在当前噪声状态下对真实反向跳跃速率的条件期望,不可约成本是前向过程 $Z_t$ 毁坏关于原始数据 $Z_0$ 的信息速率 $- frac{d}{dt}I(Z_0; Z_t)$,因此所有加噪过程共享相同的最佳可实现负ELBO——即数据熵。对于具有词元分解噪声的序列,预言投影给出三种精确坐标:去噪器、空腔(桥接插件)和得分,并可闭式转换。该框架明确了文献中各类损失实际优化的目标,恢复了MDM、UDM、SEDD和GIDD作为特例;解释了为何掩码扩散中去噪器与空腔一致而均匀扩散中不一致;证明了去噪器参数化使均匀ELBO在初始化时发散,而桥接插件仍保持有限;并在一个可解析求解的模型上数值验证了每一恒等式,无近似误差。
原文摘要 · Abstract (English)
What does a discrete diffusion model learn: a denoiser, a score ratio, or a bridge plug-in predictor? At the level of jump rates, these are one object in different coordinates, and reading a neural network in the wrong coordinate changes the process being trained and sampled. Starting with a rigorous derivation of the continuous-time Markov chain (CTMC) ELBO for any noising process, boundary terms included, we prove the \emph{Oracle Distance} theorem: the negative ELBO is exactly equal to the data entropy plus the path KL from the oracle reverse process to the learned one, not merely a bound. Its unique optimizer is therefore the conditional expectation of the true reverse jump rate given the current noisy state, and its irreducible cost is the rate at which the forward process $Z_t$ destroys information about the clean data $Z_0$, $-\tfrac{d}{dt}I(Z_0; Z_t)$, so every noising process shares the same best achievable negative ELBO: the data entropy. For sequences with token-factorizing noise, the oracle projection yields three exact coordinates for the optimizer: denoiser, cavity (bridge plug-in), and score, with closed-form conversions among them. This framework identifies which law each loss in the literature actually optimizes, recovering MDM, UDM, SEDD, and GIDD as special cases; explains why denoiser and cavity coincide for masked diffusion but not for uniform diffusion; proves that a denoiser parameterization makes the uniform ELBO diverge at initialization while the bridge plug-in stays finite; and calibrates ELBO implementations exactly at initialization. Every identity is verified numerically, without approximation, on an exactly solvable model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。