arXiv:2509.13866cs.LGcs.AI2025-09NeurIPS被引 2

将掩码扩散模型统一为能量最小化问题,提升采样效率。

Masked Diffusion Models as Energy Minimization

  • 从最优传输视角揭示掩码扩散模型的三种能量等价性。
  • 设计基于Beta分布的插值调度,低步数采样性能显著优于人工设计。
  • 无需修改模型即可高效调优,适合需要快速生成的场景。

我们提出一个系统的理论框架,将掩码扩散模型(MDMs)解释为离散最优传输中的能量最小化问题。具体而言,我们证明了三种不同的能量形式——动能、条件动能和测地线能量——在MDM结构下数学等价,且当掩码调度满足闭式最优条件时,MDM同时最小化这三种能量。这一统一不仅阐明了MDM的理论基础,还启发了采样策略的改进。通过用Beta分布参数化插值调度,我们将调度设计空间简化为可处理的二维搜索,实现无需修改模型的高效后训练调优。在合成与真实世界基准上的实验表明,基于能量设计的调度在低步数采样设置下显著优于人工设计基线。

原文摘要 · Abstract (English)

We present a systematic theoretical framework that interprets masked diffusion models (MDMs) as solutions to energy minimization problems in discrete optimal transport. Specifically, we prove that three distinct energy formulations--kinetic, conditional kinetic, and geodesic energy--are mathematically equivalent under the structure of MDMs, and that MDMs minimize all three when the mask schedule satisfies a closed-form optimality condition. This unification not only clarifies the theoretical foundations of MDMs, but also motivates practical improvements in sampling. By parameterizing interpolation schedules via Beta distributions, we reduce the schedule design space to a tractable 2D search, enabling efficient post-training tuning without model modification. Experiments on synthetic and real-world benchmarks demonstrate that our energy-inspired schedules outperform hand-crafted baselines, particularly in low-step sampling settings.

扩散模型能量最小化采样优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。