arXiv:2510.22926cs.LG2025-10被引 1

简化扩散语言模型的去噪损失,提升训练稳定性和生成效率。

Simple Denoising Diffusion Language Models

  • 仅优化被噪声替换的词元,降低复杂度。
  • 在多个数据集上达到与复杂方法相当的生成性能。
  • 适合大规模语言模型训练,兼具高效与稳定优势。

近期的均匀状态扩散模型(USDMs)从均匀先验初始化,因具备内在自纠正能力,相比掩码扩散模型可实现更快的文本生成。然而,它们仍依赖复杂的损失函数设计,带来额外计算开销,限制了可扩展性。本文提出一种简化的基于去噪的损失,仅优化被噪声替换的词元,稳定了训练过程,并在性能上达到先前复杂目标方法的水平。此外,引入一种高效的正则化项,缓解了输出分布趋向均匀的问题,进一步提升了性能。我们在广泛使用的文本数据集上对USDM模型进行预训练,验证了所提损失方法的有效性与高效性。更重要的是,结论在更大模型上依然成立,展现出大规模训练的强大潜力。

原文摘要 · Abstract (English)

Recent Uniform State Diffusion Models (USDMs), initialized from a uniform prior, offer the promise of fast text generation due to their inherent self-correction ability compared to masked diffusion models. However, they still rely on complex loss formulations with additional computational overhead, which hinders scalability. In this work, we explore a simplified denoising-based loss for USDMs that optimizes only noise-replaced tokens, stabilizing training while matching the performance of prior methods with more complex objectives. In addition, we introduce an efficient regularization term to mitigate corruption toward uniform output distributions, which further improves performance. We demonstrate the effectiveness and efficiency of our simple and improved loss formulations by pretraining models on widely used text datasets for USDMs. More importantly, our conclusions scale to larger models, showing strong potential for large-scale training.

扩散模型文本生成去噪高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。