arXiv:2510.19304cs.LG2025-10中稿 · ICLR被引 7

提出新方法突破离散扩散模型采样瓶颈,生成更连贯文本

Loopholing Discrete Diffusion: Deterministic Bypass of the Sampling Wall

  • 通过确定性潜变量路径保留分布信息,避免采样后信息丢失
  • 生成困惑度降低61%,接近甚至超越自回归模型表现
  • 适合追求高效非自回归生成的科研与工程人员

离散扩散模型通过并行解码提供了自回归生成的替代方案,但面临采样墙问题:一旦进行类别采样,丰富的分布信息会坍缩为独热向量,无法跨步骤传递,导致后续步骤仅能基于有限信息运行。为缓解此问题,我们提出Loopholing机制,通过确定性潜变量路径保留信息,构建了Loopholing离散扩散模型(LDDMs)。采用自条件化训练策略,无需展开完整去噪轨迹,实现高效训练。实验表明,LDDMs在生成困惑度上相比基线最高降低61%,显著缩小甚至超越了与自回归模型的差距,生成文本更连贯。应用于推理任务时,也在Countdown和Game of 24等算术基准上提升性能。结果还表明,Loopholing可减少冗余步骤与振荡,为高质量非自回归文本生成提供通用有效路径。

原文摘要 · Abstract (English)

Discrete diffusion models offer a promising alternative to autoregressive generation through parallel decoding, but they suffer from a sampling wall: once categorical sampling occurs, rich distributional information collapses into one-hot vectors and cannot be propagated across steps, forcing subsequent steps to operate with limited information. To mitigate this problem, we introduce Loopholing, a novel and simple mechanism that preserves this information via a deterministic latent pathway, leading to Loopholing Discrete Diffusion Models (LDDMs). Trained efficiently with a self-conditioning strategy that avoids unrolling the full denoising trajectory, LDDMs achieve substantial gains-reducing generative perplexity by up to 61% over prior baselines, thereby closing (and in some cases surpassing) the gap with autoregressive models, and producing more coherent text. Applied to reasoning tasks, LDDMs also improve performance on arithmetic benchmarks such as Countdown and Game of 24. These results also indicate that loopholing mitigates idle steps and oscillations, providing a general and effective path toward high-quality non-autoregressive text generation.

离散扩散非自回归文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。