用可学习的高斯混合分布提升扩散模型采样效果
End-To-End Learning of Gaussian Mixture Priors for Diffusion Sampler
- 直接训练高斯混合先验,替代固定高斯分布
- 在多个真实与合成任务中显著提升采样性能
- 适合需要高质量生成且避免模式崩溃的场景
通过变分推断优化的扩散模型已成为从非归一化目标分布中生成样本的有力工具。这些模型通过模拟随机微分方程生成样本,从一个简单易处理的先验(通常是高斯分布)开始。然而,当先验支持域与目标分布差异较大时,扩散模型往往难以有效探索或产生较大的离散化误差。此外,学习先验分布可能导致模式崩溃,这由反向KL散度的模式导向性加剧。为此,我们提出端到端可学习的高斯混合先验(GMPs)。GMPs 提供更优的探索控制、对目标支持域的自适应能力以及更强的表达力以缓解模式崩溃。我们进一步利用混合模型结构,提出一种在训练过程中迭代添加混合成分的策略。实验结果表明,在无需额外目标评估的前提下,使用 GMPs 在一系列真实世界和合成基准问题上均实现显著性能提升。
原文摘要 · Abstract (English)
Diffusion models optimized via variational inference (VI) have emerged as a promising tool for generating samples from unnormalized target densities. These models create samples by simulating a stochastic differential equation, starting from a simple, tractable prior, typically a Gaussian distribution. However, when the support of this prior differs greatly from that of the target distribution, diffusion models often struggle to explore effectively or suffer from large discretization errors. Moreover, learning the prior distribution can lead to mode-collapse, exacerbated by the mode-seeking nature of reverse Kullback-Leibler divergence commonly used in VI. To address these challenges, we propose end-to-end learnable Gaussian mixture priors (GMPs). GMPs offer improved control over exploration, adaptability to target support, and increased expressiveness to counteract mode collapse. We further leverage the structure of mixture models by proposing a strategy to iteratively refine the model by adding mixture components during training. Our experimental results demonstrate significant performance improvements across a diverse range of real-world and synthetic benchmark problems when using GMPs without requiring additional target evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。