arXiv:2510.13117cs.LGcs.AI2025-10被引 8

揭示了文本掩码扩散模型的推理能力边界与效率优势。

On the Reasoning Abilities of Masked Diffusion Language Models

  • 将掩码扩散模型与思维链框架关联,建立理论等价性。
  • 证明其可解决所有思维链增强模型能解的问题,且部分任务更快。
  • 特别适合并行计算加速的正则语言等推理任务,适合追求效率的研究者。

文本掩码扩散模型(MDMs)为传统自回归语言模型提供了一种高效替代方案,其并行生成机制提升了计算效率,但其推理能力及并行性带来的局限性仍不明确。本文在有限精度对数宽度设定下,将MDMs与思维链(CoT)和填充循环变换器(PLTs)框架关联,证明了MDMs与多项式填充的PLTs在此设定下等价,并且能够求解所有思维链增强的Transformer所能解决的问题。此外,我们展示了若干类问题(包括正则语言),在这些任务中,由于并行生成的优势,MDMs的推理效率显著高于思维链Transformer。

原文摘要 · Abstract (English)

Masked diffusion models (MDMs) for text offer a compelling alternative to traditional autoregressive language models. Parallel generation makes them efficient, but their computational capabilities and the limitations inherent in their parallelism remain largely unexplored. To this end, we characterize what types of reasoning problems MDMs can provably solve and how efficiently. We do this by connecting MDMs to the well-understood reasoning frameworks of chain of thought (CoT) and padded looped transformers (PLTs) in the finite-precision log-width setting: We show that MDMs and polynomially-padded PLTs are, in fact, equivalent in this setting, and that MDMs can solve all problems that CoT-augmented transformers can. Moreover, we showcase classes of problems (including regular languages) for which MDMs are inherently more efficient than CoT transformers, where parallel generation allows for substantially faster reasoning.

扩散模型推理能力并行生成语言建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。