arXiv:2507.04341stat.MLcs.AI2025-07ICLR被引 7

改进离散扩散语言模型的困惑度上界与训练效率。

Efficient Perplexity Bound and Ratio Matching in Discrete Diffusion Language Models

  • 基于连续时间马尔可夫链推导出新的KL散度定理,优化困惑度上界。
  • 用去噪交叉熵替代得分熵,使困惑度降低10%,训练提速15%。
  • 提出新型转移率矩阵并解析其指数形式,支持高效生成与训练。

尽管连续扩散模型在建模连续分布方面表现优异,但在分类数据上的应用效果不佳。近期研究发现,在连续时间离散马尔可夫链(CTMC)框架下通过得分熵进行比率匹配,可作为语言建模中自回归模型的有力替代方案。为增强该框架,我们首先提出三个关于数据分布与学习分布之间KL散度的新定理,其结果是连续扩散模型相关理论在离散场景下的对应版本,从而推导出更优的困惑度上界。其次,实证表明,通过最小化干净数据与噪声数据间的去噪交叉熵进行比率匹配,可使模型在困惑度/生成困惑度上比使用得分熵的模型降低最多10%,且训练速度提升15%。为进一步验证该结论,我们引入并评估了一种新型的CTMC转移率矩阵,支持预测精炼,并推导出其矩阵指数的解析表达式,从而实现条件比率的高效计算,促进训练与生成过程的加速。

原文摘要 · Abstract (English)

While continuous diffusion models excel in modeling continuous distributions, their application to categorical data has been less effective. Recent work has shown that ratio-matching through score-entropy within a continuous-time discrete Markov chain (CTMC) framework serves as a competitive alternative to autoregressive models in language modeling. To enhance this framework, we first introduce three new theorems concerning the KL divergence between the data and learned distribution. Our results serve as the discrete counterpart to those established for continuous diffusion models and allow us to derive an improved upper bound of the perplexity. Second, we empirically show that ratio-matching performed by minimizing the denoising cross-entropy between the clean and corrupted data enables models to outperform those utilizing score-entropy with up to 10% lower perplexity/generative-perplexity, and 15% faster training steps. To further support our findings, we introduce and evaluate a novel CTMC transition-rate matrix that allows prediction refinement, and derive the analytic expression for its matrix exponential which facilitates the computation of conditional ratios thus enabling efficient training and generation.

离散扩散语言建模比率匹配高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。