arXiv:2502.09622cs.LGcs.AI2025-02NeurIPS被引 52

揭示扩散语言模型在不同评估指标下的效率与精度权衡

Theoretical Benefit and Limitation of Diffusion Language Model

  • 从理论上分析掩码扩散模型的生成机制与收敛特性
  • 用困惑度评估时可并行生成且性能接近最优,但用序列错误率则需线性增加采样步数
  • 为理解扩散模型优势与局限提供首个理论框架,适合研究生成模型的学者

扩散语言模型作为文本生成的新方法备受关注。尽管每步可并行采样多个词元,理论上应比自回归模型更高效,但其效率-精度权衡尚不明确。本文对广泛使用的掩码扩散模型(MDM)进行严格理论分析发现,其效果高度依赖评估指标。在温和条件下,当使用困惑度作为评价标准时,无论序列长度如何,MDM均能在固定采样步数内达到近似最优困惑度,表明效率与性能可兼得。然而,若采用序列错误率(如推理链正确性指标),则所需采样步数必须随序列长度线性增长,从而丧失并行优势。该分析建立了首个关于MDM优劣的理论基础,所有结论均经实证验证。

原文摘要 · Abstract (English)

Diffusion language models have emerged as a promising approach for text generation. One would naturally expect this method to be an efficient replacement for autoregressive models since multiple tokens can be sampled in parallel during each diffusion step. However, its efficiency-accuracy trade-off is not yet well understood. In this paper, we present a rigorous theoretical analysis of a widely used type of diffusion language model, the Masked Diffusion Model (MDM), and find that its effectiveness heavily depends on the target evaluation metric. Under mild conditions, we prove that when using perplexity as the metric, MDMs can achieve near-optimal perplexity in sampling steps regardless of sequence length, demonstrating that efficiency can be achieved without sacrificing performance. However, when using the sequence error rate--which is important for understanding the "correctness" of a sequence, such as a reasoning chain--we show that the required sampling steps must scale linearly with sequence length to obtain "correct" sequences, thereby eliminating MDM's efficiency advantage over autoregressive models. Our analysis establishes the first theoretical foundation for understanding the benefits and limitations of MDMs. All theoretical findings are supported by empirical studies.

扩散模型语言建模理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。