提出新方法提升离散扩散模型采样稳定性与生成质量。
Mean-to-Score Discrete Diffusion: Posterior-Mean Denoisers for Score Entropy

- 用后验均值映射生成得分,确保符合贝叶斯可实现性。
- 在CIFAR-10上将测试BPD从3.173降至3.129,FID降低至3.129。
- 适用于多种噪声机制,170M模型在128步下生成PPL达143.3。
Score Entropy Discrete Diffusion (SEDD) 通过无约束的正得分比参数化离散反向过程,但仅保证非负跳跃率,不确保贝叶斯可实现性:噪声状态下的得分比未必由前向核诱导的清洁标记后验共同生成。尽管得分熵损失具有正确的总体最优解,却未在远离该解时强制此约束。在训练后的纯均匀SEDD检查点中,约四分之一的完整得分向量违反坐标框约束,超过一半虽在框内仍与任何有效后验不兼容,可能导致有限步采样中的负预归一化权重。将原始得分投影至桥多面体可消除所有观测到的负权重,并在不改变采样器的前提下将外部生成PPL从203.6提升至175.1。本文引入【均值到得分】(M2S)方法,预测清洁标记后验均值,并通过精确的核相关线性映射转换为得分。该构造适用于任意满足弱支撑条件的坐标式连续时间马尔可夫链(CTMC)。对于均匀污染,其将概率单纯形映射至桥多面体;对于吸收掩码污染,所得目标恰好恢复MD4。在2840万参数的CIFAR-10控制对比中,M2S将测试BPD从3.173降至3.129,FID-50k从\CifarSEDDFID降至\CifarMtwoSFID。一个在约2620亿个OpenWebText token上训练的1.7亿参数M2S模型,在所有采样预算下均优于评估的纯均匀SEDD、GIDD和神经CTMC检查点,128步下达到生成PPL 143.3,优于最强基线的183.6。
原文摘要 · Abstract (English)
Score Entropy Discrete Diffusion (SEDD) parameterizes discrete reverse processes with unconstrained positive score ratios. While positivity guarantees nonnegative reverse jump rates, it does not ensure Bayes realizability: ratios at a noisy state need not be jointly induced by any clean-token posterior under the forward kernel. The score-entropy loss has the correct population optimum but does not enforce this constraint away from it. In a trained pure-uniform SEDD checkpoint, roughly one quarter of complete score vectors violate the coordinate box, while more than half lie inside it yet remain materially incompatible with any valid posterior. Such violations can produce negative pre-normalization weights in finite-step sampling. Projecting raw scores onto the bridge polytope removes all observed negative weights and improves external generative PPL from $203.6$ to $175.1$ without changing the sampler. We introduce \emph{mean-to-score} (M2S), which predicts a clean-token posterior mean and converts it to the score through an exact kernel-dependent linear map. The construction applies to any known coordinate-wise continuous-time Markov chain (CTMC) satisfying a mild support condition. For uniform corruption, it maps the probability simplex onto the bridge polytope; for absorbing-mask corruption, the resulting objective recovers MD4 exactly. In a controlled 28.4M-parameter CIFAR-10 comparison, M2S lowers test BPD from $3.173$ to $3.129$ and FID-50k from $\CifarSEDDFID$ to $\CifarMtwoSFID$. A 170M-parameter M2S model trained on about 262B OpenWebText token slots outperforms the evaluated pure-uniform SEDD, GIDD, and Neural CTMC checkpoints at every tested sampling budget, reaching generative PPL $143.3$ at 128 steps versus $183.6$ for the strongest pure-uniform baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。