提出隐空间随机插值,实现端到端生成模型联合训练。
Latent Stochastic Interpolants
- 在隐空间中构建连续时间的变分推断目标函数
- 在ImageNet上实现高质量图像生成,超越传统扩散模型
- 适合需要灵活生成与高效训练的生成模型研究者
随机插值(SI)是一种强大的生成建模框架,可灵活地在两个概率分布间进行转换。然而,其在联合优化的隐变量模型中的应用尚未被探索,因为该方法需要直接访问两个分布的样本。本文提出隐空间随机插值(LSI),实现了编码器、解码器与隐空间SI模型的端到端联合学习。通过在连续时间下推导出一个严谨的证据下界(ELBO)目标函数,LSI能够在隐空间中学习有效表示,并将任意先验分布转化为由编码器定义的聚合后验分布。该方法规避了传统扩散模型的简单先验假设,同时避免了在高维观测空间直接应用SI带来的计算开销,保留了SI框架的生成灵活性。我们在标准的大规模ImageNet生成基准上进行了全面实验,验证了LSI的有效性。
原文摘要 · Abstract (English)
Stochastic Interpolants (SI) is a powerful framework for generative modeling, capable of flexibly transforming between two probability distributions. However, its use in jointly optimized latent variable models remains unexplored as it requires direct access to the samples from the two distributions. This work presents Latent Stochastic Interpolants (LSI) enabling joint learning in a latent space with end-to-end optimized encoder, decoder and latent SI models. We achieve this by developing a principled Evidence Lower Bound (ELBO) objective derived directly in continuous time. The joint optimization allows LSI to learn effective latent representations along with a generative process that transforms an arbitrary prior distribution into the encoder-defined aggregated posterior. LSI sidesteps the simple priors of the normal diffusion models and mitigates the computational demands of applying SI directly in high-dimensional observation spaces, while preserving the generative flexibility of the SI framework. We demonstrate the efficacy of LSI through comprehensive experiments on the standard large scale ImageNet generation benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。