arXiv:2602.15014cs.LGcs.CL2026-02被引 20

挑战掩码扩散模型的统治地位,发现更高效采样方法更具潜力

Scaling Beyond Masked Diffusion Language Models

  • 对比统一状态与插值离散扩散方法的缩放规律
  • 1.7B参数下统一状态模型在GSM8K上超越自回归和掩码扩散模型
  • 采样效率比掩码扩散高12%,适合追求生成速度的场景

扩散语言模型因其潜在的快速生成能力,成为自回归模型的有力替代。在离散扩散方法中,掩码扩散目前占主导地位,主要因其在语言建模基准上表现优异的困惑度。本文首次对统一状态和插值离散扩散方法进行缩放定律研究,并发现通过简单的交叉熵目标训练,掩码扩散模型可实现约12%的FLOPs效率提升。结果表明,困惑度在同类方法内具有参考价值,但跨类比较时可能误导——某些似然性能较差但采样更快的方法在速度-质量帕累托前沿上更优。将所有方法扩展至1.7B参数后,统一状态扩散模型在似然基准上仍具竞争力,且在GSM8K任务上优于自回归与掩码扩散模型,尽管其验证困惑度更低。项目页面提供代码、模型检查点及视频教程:http://s-sahoo.github.io/scaling-dllms

原文摘要 · Abstract (English)

Diffusion language models are a promising alternative to autoregressive models due to their potential for faster generation. Among discrete diffusion approaches, Masked diffusion currently dominates, largely driven by strong perplexity on language modeling benchmarks. In this work, we present the first scaling law study of uniform-state and interpolating discrete diffusion methods. We also show that Masked diffusion models can be made approximately 12% more FLOPs-efficient when trained with a simple cross-entropy objective. We find that perplexity is informative within a diffusion family but can be misleading across families, where models with worse likelihood scaling may be preferable due to faster and more practical sampling, as reflected by the speed-quality Pareto frontier. These results challenge the view that Masked diffusion is categorically the future of diffusion language modeling and that perplexity alone suffices for cross-algorithm comparison. Scaling all methods to 1.7B parameters, we show that uniform-state diffusion remains competitive on likelihood-based benchmarks and outperforms autoregressive and Masked diffusion models on GSM8K, despite worse validation perplexity. We provide the code, model checkpoints, and video tutorials on the project page: http://s-sahoo.github.io/scaling-dllms

扩散模型语言建模生成效率缩放定律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。