arXiv:2512.22630cs.CL2025-12被引 4

探索扩散模型生成文本时离散性的作用,揭示现有方法的结构性缺陷。

On the Role of Discreteness in Diffusion LLMs

  • 区分嵌入空间连续扩散与词元空间离散扩散两类方法
  • 发现统一扰动和词元独立训练会破坏文本结构与依赖关系
  • 为更符合语言特性的扩散生成模型提供设计方向

扩散模型在语言生成中具有并行解码和迭代优化等优势,但文本的离散性和高结构特性使其难以直接应用。本文从扩散过程与语言建模两个视角出发,梳理出五项区分扩散机制与语言需求的核心属性。首先将现有方法分为嵌入空间的连续扩散与词元空间的离散扩散。分析表明,二者仅满足部分关键属性,存在结构性权衡。通过对近期大型扩散语言模型的考察,识别出两大核心问题:(i) 统一扰动不考虑信息在位置间的分布差异;(ii) 词元级边缘训练无法捕捉并行解码中的多词元依赖关系。这些发现推动未来研究向更贴近文本结构的扩散过程发展。

原文摘要 · Abstract (English)

Diffusion models offer appealing properties for language generation, such as parallel decoding and iterative refinement, but the discrete and highly structured nature of text challenges the direct application of diffusion principles. In this paper, we revisit diffusion language modeling from the view of diffusion process and language modeling, and outline five properties that separate diffusion mechanics from language-specific requirements. We first categorize existing approaches into continuous diffusion in embedding space and discrete diffusion over tokens. We then show that each satisfies only part of the five essential properties and therefore reflects a structural trade-off. Through analyses of recent large diffusion language models, we identify two central issues: (i) uniform corruption does not respect how information is distributed across positions, and (ii) token-wise marginal training cannot capture multi-token dependencies during parallel decoding. These observations motivate diffusion processes that align more closely with the structure of text, and encourage future work toward more coherent diffusion language models.

扩散模型语言生成离散性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。