通过分阶段扰动提升扩散语言模型生成多样性
Time-Annealed Perturbation Sampling: Diverse Generation for Diffusion Language Models
- 早期引入扰动实现语义分支,后期减少扰动保证流畅性
- 在创意写作与推理任务中提升多样性,不降低生成质量
- 无需训练,适配多种扩散语言模型架构
扩散语言模型(Diffusion-LMs)为文本生成引入了显式的时序结构,但如何利用这一结构控制生成多样性以探索多种有效语义或推理路径仍不明确。本文发现,扩散语言模型与图像生成中的扩散模型类似,存在时间分工:早期去噪步骤决定全局语义结构,后期则聚焦局部词汇优化。基于此,我们提出无需训练的推理策略Time-Annealed Perturbation Sampling(TAPS),在扩散过程早期鼓励语义分支,同时逐步减少扰动以保持流畅性与指令遵循。TAPS兼容非自回归与半自回归扩散骨架,在LLaDA和TraDo上验证,显著提升创意写作与推理基准下的输出多样性,且不损害生成质量。
原文摘要 · Abstract (English)
Diffusion language models (Diffusion-LMs) introduce an explicit temporal dimension into text generation, yet how this structure can be leveraged to control generation diversity for exploring multiple valid semantic or reasoning paths remains underexplored. In this paper, we show that Diffusion-LMs, like diffusion models in image generation, exhibit a temporal division of labor: early denoising steps largely determine the global semantic structure, while later steps focus on local lexical refinement. Building on this insight, we propose Time-Annealed Perturbation Sampling (TAPS), a training-free inference strategy that encourages semantic branching early in the diffusion process while progressively reducing perturbations to preserve fluency and instruction adherence. TAPS is compatible with both non-autoregressive and semi-autoregressive Diffusion backbones, demonstrated on LLaDA and TraDo in our paper, and consistently improves output diversity across creative writing and reasoning benchmarks without compromising generation quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。