重加权损失可提升扩散模型的变分下界,显著改善生成质量。
Demystifying Diffusion Objectives: Reweighted Losses are Better Variational Bounds
- 构建时序变分下界级联,改进标准证据下界。
- 在像素空间图像建模中性能逼近连续扩散模型。
- 为掩码图像模型的简单加权方案提供理论支持。
我们推导了广泛用于训练扩散模型的重加权损失的新理论解释。方法基于构建数据对数似然的时序依赖变分下界级联,可严格优于标准证据下界,并降低数据-模型之间的KL散度。结合这些下界得到的重加权目标适用于任意生成式扩散模型,包括连续高斯扩散和掩码(离散)扩散模型。我们在掩码扩散模型中验证该框架,报告在像素空间图像建模中相比以往训练损失有显著提升,生成样本质量接近连续扩散模型。结果还为掩码图像模型中广泛使用的简单加权方案提供了理论依据。
原文摘要 · Abstract (English)
We derive a new theoretical interpretation of the reweighted losses that are widely used for training diffusion models. Our method is based on constructing a cascade of time-dependent variational lower bounds on the data log-likelihood, that provably improves upon the standard evidence lower bound and results in reduced data-model KL-divergences. Combining such bounds gives rise to reweighted objectives that can be applied to any generative diffusion model including both continuous Gaussian diffusion and masked (discrete) diffusion models. Then, we showcase this framework in masked diffusion and report significant improvements over previous training losses in pixel-space image modeling, approaching sample quality comparable to continuous diffusion models. Our results also provide a theoretical justification for the simple weighting scheme widely used in masked image models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。