arXiv:2510.03289cs.LGcs.AI2025-10被引 2

揭示掩码扩散模型无法实现并行生成的根本原因

Why mask diffusion does not work

  • 从吸收扩散机制出发,分析其内在限制
  • 实证表明该模型难以支持真正的并行生成
  • 提出有效训练与推理策略,适合研究者参考

扩散语言模型相较于自回归模型的主要优势在于支持并行生成和双向注意力,从而实现更可控的生成过程。近年来,开源的掩码扩散语言模型涌现,多数基于一种称为吸收扩散的变体。然而,本文揭示了掩码扩散在实现并行生成和双向注意力方面存在固有困难。通过理论与实验分析,我们进一步提出了适用于掩码扩散的有效训练与推理策略。

原文摘要 · Abstract (English)

The main advantages of diffusion language models over autoregressive (AR) models lie in their ability to support parallel generation and bidirectional attention, enabling a more controllable generation process. In recent years, open-source mask diffusion language models have emerged, most of which are based on a variant known as absorbing diffusion. However, this paper demonstrates why mask diffusion faces inherent difficulties in achieving parallel generation and bidirectional attention. We also propose the most effective training and inference strategies for mask diffusion.

扩散模型语言生成模型机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。