arXiv:2506.08316cs.LGstat.ML2025-06NeurIPS被引 16

揭示掩码扩散为何更优,提出可通用的跳跃调度条件模型

Why Masking Diffusion Works: Condition on the Jump Schedule for Improved Discrete Diffusion

  • 将跳跃时间分布显式建模于扩散过程,替代渐进去噪
  • 新模型在图像、文本、蛋白质数据上均超越掩码扩散
  • 适合需高效生成且有先验知识的数据类型

离散扩散模型通过马尔可夫过程逐步去除噪声生成高质量样本。理论上,这种渐进生成具有诸多优势,如可融入归纳偏置和使用更优采样算法。然而实践中,性能最佳的模型却是不进行渐进去噪的掩码扩散。本文指出其优越性源于连续与离散马尔可夫过程的根本差异:离散过程以固定速率发生突变跳跃,而掩码扩散利用了已知的跳跃时间分布,仅学习跳转目标。我们进一步将此思想推广至任意离散扩散模型,提出调度条件离散扩散(SCUD),使其能显式嵌入跳跃时间分布。在包含图像、文本和蛋白质数据归纳偏置的无噪过程上应用SCUD,所构建模型均优于掩码扩散。

原文摘要 · Abstract (English)

Discrete diffusion models, like continuous diffusion models, generate high-quality samples by gradually undoing noise applied to datapoints with a Markov process. Gradual generation in theory comes with many conceptual benefits; for example, inductive biases can be incorporated into the noising Markov process, and access to improved sampling algorithms. In practice, however, the consistently best performing discrete diffusion model is, surprisingly, masking diffusion, which does not denoise gradually. Here we explain the superior performance of masking diffusion by noting that it makes use of a fundamental difference between continuous and discrete Markov processes: discrete Markov processes evolve by discontinuous jumps at a fixed rate and, unlike other discrete diffusion models, masking diffusion builds in the known distribution of jump times and only learns where to jump to. We show that we can similarly bake in the known distribution of jump times into any discrete diffusion model. The resulting models - schedule-conditioned discrete diffusion (SCUD) - generalize classical discrete diffusion and masking diffusion. By applying SCUD to models with noising processes that incorporate inductive biases on images, text, and protein data, we build models that outperform masking.

扩散模型离散生成跳跃调度掩码扩散

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。