arXiv:2511.08946cs.LG2025-11

用非保体积变换改进条件变分自编码器,生成更清晰多样图像

Improving Conditional VAE with Non-Volume Preserving transformations

  • 引入非保体积变换建模条件隐空间分布,突破传统假设
  • 相比旧方法,FID降低4%,对数似然提升7.6%
  • 适合想提升生成质量的图像生成研究者

变分自编码器与生成对抗网络曾是生成模型的主流,但2022年后被基于扩散的模型超越。受此影响,传统模型改进趋于停滞。本文以经典方式探索条件变分自编码器(CVAE)在图像生成中的应用,旨在将特定属性融入生成图像。传统VAE常生成模糊且多样性不足的图像,本文通过将高斯解码器方差设为可学习参数来缓解该问题。此前的CVAE研究假设给定标签的隐空间条件分布等于先验分布,这在现实中并不成立。本文提出使用非保体积(NVP)变换估计该条件分布,显著提升生成效果:相较先前方法,FID降低4%,对数似然提高7.6%。

原文摘要 · Abstract (English)

Variational Autoencoders and Generative Adversarial Networks remained the state-of-the-art (SOTA) generative models until 2022. Now they are superseded by diffusion-based models. Efforts to improve traditional models have stagnated as a result. In old-school fashion, we explore image generation with conditional Variational Autoencoders (CVAE) to incorporate desired attributes within the images. VAEs are known to produce blurry images with less diversity; we refer to a method that solves this issue by leveraging the variance of the gaussian decoder as a learnable parameter during training. Previous works on CVAEs assumed that the conditional distribution of the latent space given the labels is equal to the prior distribution, which is not the case in reality. We show that estimating it using Non-Volume Preserving (NVP) transformations results in better image generation than existing methods by reducing the FID by 4% and increasing log likelihood by 7.6% compared to the previous cases.

变分自编码器图像生成生成模型条件生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。