arXiv:2605.18530cs.CLcs.AI2026-05被引 7

连续扩散模型经优化后可媲美离散模型,生成质量与效率俱佳。

Continuous Diffusion Scales Competitively with Discrete Diffusion for Language

论文配图:Continuous Diffusion Scales Competitively with Discrete Diffusion for Language
图 1 · 摘自论文原文
  • 通过统一架构对比,重构连续扩散语言模型RePlaid
  • 在OpenWebText上达22.1的最低困惑度,优于现有连续模型
  • 理论揭示似然训练优势:噪声调度均衡去噪难度

尽管扩散模型在语言建模领域备受关注,连续扩散模型此前被认为不如离散方法可扩展。为挑战这一观点,我们重新审视基于似然的连续扩散语言模型Plaid,构建了与现代离散模型对齐的RePlaid。在统一设置下,首次建立连续扩散语言模型的缩放定律:RePlaid的计算开销仅是自回归模型的20倍,参数更少却优于Duo,过训练阶段超越MDLM。在OpenWebText上,RePlaid实现连续模型中新的最优困惑度22.1,生成质量更优。结果表明,基于似然训练的连续扩散模型是极具竞争力且可扩展的替代方案。此外,我们提供理论分析:优化噪声调度以最小化ELBO方差,自然导致线性交叉熵(信息损失)随时间变化,均匀分布去噪难度,无需特定时间重参数化;同时发现,通过似然优化嵌入能形成结构化几何,带来最大似然提升。

原文摘要 · Abstract (English)

While diffusion has drawn considerable recent attention from the language modeling community, continuous diffusion has appeared less scalable than discrete approaches. To challenge this belief we revisit Plaid, a likelihood-based continuous diffusion language model (DLM), and construct RePlaid by aligning the architecture of Plaid with modern discrete DLMs. In this unified setting, we establish the first scaling law for continuous DLMs that rivals discrete DLMs: RePlaid exhibits a compute gap of only $20\times$ compared to autoregressive models, outperforms Duo while using fewer parameters, and outperforms MDLM in the over-trained regime. We benchmark RePlaid against recent continuous DLMs: on OpenWebText, RePlaid achieves a new state-of-the-art PPL bound of $22.1$ among continuous DLMs and superior generation quality. These results suggest that continuous diffusion, when trained via likelihood, is a highly competitive and scalable alternative to discrete DLMs. Moreover, we offer theoretical insights to understand the advantage of likelihood-based training. We show that optimizing the noise schedule to minimize the ELBO's variance naturally yields linear cross-entropy (information loss) over time. This evenly distributes denoising difficulty without any case-specific time reparameterization. In addition, we find that optimizing embeddings via likelihood creates structured geometries and drives the most significant likelihood gain.

扩散模型语言建模连续扩散似然训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。