提出联合连续与离散扩散的新型语言模型,提升推理能力与生成质量。
Coevolutionary Continuous Discrete Diffusion: Make Your Diffusion Language Model a Latent Reasoner
- 设计联合连续与离散空间的协同扩散机制,统一建模隐变量与词元。
- 在多个真实任务上实现优于离散扩散模型的生成质量与训练稳定性。
- 适合追求高表达力与良好采样效果的生成模型研究者使用。
扩散语言模型,尤其是掩码离散扩散模型,近期取得了显著进展。尽管理论和初步实证表明,通过循环变换器或连续思维链进行潜在推理具有优势,但连续扩散模型通常表现不如其离散版本。本文认为,扩散语言模型无需局限于离散空间。我们证明,连续扩散模型比离散扩散和循环变换器具有更强的表达能力。性能差距源于实际可训练性:连续扩散虽提供中间监督,但将连续表示解码回离散词元空间存在额外困难。为此,我们提出共进化连续-离散扩散(CCDD),在连续表示空间与离散词元空间的并集中定义联合多模态扩散过程,用单一模型同时在联合空间去噪。通过融合两种模态,CCDD兼具潜空间丰富的语义表达、良好的可训练性及高质量样本生成。我们还提出了有效的架构与先进训练/采样技术,在真实世界任务的语言建模实验中展现出强劲的实证性能。
原文摘要 · Abstract (English)
Diffusion language models, especially masked discrete diffusion models, have achieved great success recently. While there are some theoretical and primary empirical results showing the advantages of latent reasoning with looped transformers or continuous chain-of-thoughts, continuous diffusion models typically underperform their discrete counterparts. In this paper, we argue that diffusion language models do not necessarily need to be in the discrete space. In particular, we prove that continuous diffusion models have stronger expressivity than discrete diffusions and looped transformers. We attribute the contradiction between the theoretical expressiveness and empirical performance to their practical trainability: while continuous diffusion provides intermediate supervision that looped transformers lack, they introduce additional difficulty decoding tokens into the discrete token space from the continuous representation space. We therefore propose Coevolutionary Continuous Discrete Diffusion (CCDD), which defines a joint multimodal diffusion process on the union of a continuous representation space and a discrete token space, leveraging a single model to simultaneously denoise in the joint space. By combining two modalities, CCDD is expressive with rich semantics in the latent space, as well as good trainability and sample quality with the help of explicit discrete tokens. We also propose effective architectures and advanced training/sampling techniques for CCDD, which reveals strong empirical performance in extensive language modeling experiments on real-world tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。