arXiv:2501.09876math.NAcs.LG2025-01被引 6

提出几何保真编码器,提升生成模型训练效率与稳定性。

Geometry-Preserving Encoder/Decoder in Latent Generative Models

  • 设计新编码器/解码器框架,保持数据分布几何结构。
  • 理论证明编码器训练收敛,解码器收敛速度更快。
  • 适合追求高效稳定生成模型的研究者与工程师。

生成建模旨在生成与给定数据集相似的新样本。使用扩散模型时,主要挑战在于高维输入空间的求解。近期方法通过编码器将数据映射到低维潜在空间,在潜在空间中求解扩散模型,从而提升训练效率并达到顶尖性能。变分自编码器(VAE)是该领域最常用的编码器/解码器框架,具备学习潜在表示和生成数据的能力。本文提出一种新型编码器/解码器框架,其具有与VAE不同的理论性质,专为保持数据分布的几何结构而设计。我们展示了该几何保真编码器在编码器与解码器训练过程中的显著优势,并提供了理论结果:证明了编码器训练的收敛性,以及使用该编码器时解码器训练收敛速度更快。

原文摘要 · Abstract (English)

Generative modeling aims to generate new data samples that resemble a given dataset. When using diffusion models for this task, one of the main challenges is solving the problem in the input space, which tends to be very high-dimensional. To address this, recent approaches solve diffusion models in the latent space through an encoder that maps from the data space to a lower-dimensional latent space, improving training efficiency and achieving state-of-the-art results. The variational autoencoder (VAE) is the most commonly used encoder/decoder framework in this domain, known for its ability to learn latent representations and generate data samples. In this paper, we introduce a novel encoder/decoder framework with theoretical properties distinct from those of the VAE, specifically designed to preserve the geometric structure of the data distribution. We demonstrate the significant advantages of this geometry-preserving encoder in the training process of both the encoder and decoder. Additionally, we provide theoretical results proving convergence of the training process, including convergence guarantees for encoder training, and results showing faster convergence of decoder training when using the geometry-preserving encoder.

生成模型扩散模型几何保真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。