让扩散采样器自适应调整破坏过程,提升生成速度与质量
Adaptive Destruction Processes for Diffusion Samplers
- 解耦生成与破坏的方差,两者均以自由参数的高斯分布学习
- 在有限步数下训练双过程,收敛更快且生成质量更优
- 适合需要快速生成的场景,如条件图像生成中的潜空间采样
本文探讨了在扩散采样器中引入可训练破坏过程的挑战与优势——这类基于扩散的生成模型在无法访问真实数据样本的情况下,需对非归一化密度进行采样。不同于多数研究将扩散采样器视为连续时间模型的近似,本文将其视为离散时间策略,目标是在极少数生成步骤内产出样本。为此,我们牺牲部分理论优雅性,换取生成与破坏策略定义上的灵活性。具体地,解耦生成与破坏的方差,使两个转移核均可作为无约束高斯分布进行学习。实验表明,在生成步数受限时,同时训练生成与破坏过程能实现更快收敛和更优采样质量。通过稳健的消融实验,我们分析了稳定训练所需的设计选择。最后,我们在生成对抗网络(GAN)潜空间采样任务上验证了该方法的可扩展性,用于条件图像生成。
原文摘要 · Abstract (English)
This paper explores the challenges and benefits of a trainable destruction process in diffusion samplers -- diffusion-based generative models trained to sample an unnormalised density without access to data samples. Contrary to the majority of work that views diffusion samplers as approximations to an underlying continuous-time model, we view diffusion models as discrete-time policies trained to produce samples in very few generation steps. We propose to trade some of the elegance of the underlying theory for flexibility in the definition of the generative and destruction policies. In particular, we decouple the generation and destruction variances, enabling both transition kernels to be learned as unconstrained Gaussian densities. We show that, when the number of steps is limited, training both generation and destruction processes results in faster convergence and improved sampling quality on various benchmarks. Through a robust ablation study, we investigate the design choices necessary to facilitate stable training. Finally, we show the scalability of our approach through experiments on GAN latent space sampling for conditional image generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。