提升少步扩散语言模型生成质量,实现400倍加速
FastDiSS: Few-step Match Many-step Diffusion Language Model on Sequence-to-Sequence Generation--Full Version
- 通过扰动自条件信号匹配推理噪声,增强少步采样鲁棒性
- 在多个基准上实现比标准扩散模型更优的生成质量
- 适合追求高速生成且对精度要求高的实际应用
自条件机制是连续扩散语言模型成功的关键,可修正先前错误。但在最需要快速推理的少步采样场景下,其性能显著下降。本研究发现,当仅使用少量去噪步骤时,不准确的自条件会引入显著近似误差,且该误差在去噪过程中累积,最终主导样本质量。为此,我们提出一种新型训练框架:在训练中扰动自条件信号以匹配推理噪声,提升对前期估计误差的鲁棒性。此外,引入基于词元的噪声感知机制,防止训练饱和,改善优化过程。在多个条件生成基准上的实验证明,该方法优于标准连续扩散模型,推理速度最快可达400倍提升,且在与单步扩散框架的对比中仍具竞争力。
原文摘要 · Abstract (English)
Self-conditioning has been central to the success of continuous diffusion language models, as it allows models to correct previous errors. Yet its ability degrades precisely in the regime where diffusion is most attractive for deployment: few-step sampling for fast inference. In this study, we show that when models only have a few denoising steps, inaccurate self-conditioning induces a substantial approximation gap; this mistake compounds across denoising steps and ultimately dominate the sample quality. To address this, we propose a novel training framework that handles these errors during learning by perturbing the self-conditioning signal to match inference noise, improving robustness to prior estimation errors. In addition, we introduce a token-level noise-awareness mechanism that prevents training from saturation, hence improving optimization. Extensive experiments across conditional generation benchmarks demonstrate that our framework surpasses standard continuous diffusion models while providing up to 400x faster inference speed, and remains competitive against other one-step diffusion frameworks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。