用多项式层次结构改进VAE,提升生成质量与表征可解释性。
PH-VAE: A Polynomial Hierarchical Variational Autoencoder Towards Disentangled Representation Learning
- 构建多项式层次结构,替代传统VAE的编码方式。
- 新提出的多项式散度使重建图像更清晰、分布更准确。
- 具备解耦表征能力,适合复杂数据生成任务。
变分自编码器(VAE)是一种高效建模复杂数据分布(如图像、文本)的生成方法,但存在潜在缺陷:隐变量不可解释、超参数难调、生成结果模糊、损失函数设计导致信息丢失、过拟合及小样本下的源点引力效应等,影响复杂分布数据的生成效果。本文提出并开发了多项式层次变分自编码器(PH-VAE),采用多项式层次数据格式进行分布生成与重构。同时,在损失函数中引入新型多项式散度以替代或泛化经典的KL散度,显著提升了重构分布的准确性和可复现性,以及重构图像的质量,且在相同数据集规模下捕捉到更高分辨率的细节。此外,实验表明所提PH-VAE具备一定的解耦表征学习能力。
原文摘要 · Abstract (English)
The variational autoencoder (VAE) is a simple and efficient generative artificial intelligence method for modeling complex probability distributions of various types of data, such as images and texts. However, it suffers some main shortcomings, such as lack of interpretability in the latent variables, difficulties in tuning hyperparameters while training, producing blurry, unrealistic downstream outputs or loss of information due to how it calculates loss functions and recovers data distributions, overfitting, and origin gravity effect for small data sets, among other issues. These and other limitations have caused unsatisfactory generation effects for the data with complex distributions. In this work, we proposed and developed a polynomial hierarchical variational autoencoder (PH-VAE), in which we used a polynomial hierarchical date format to generate or to reconstruct the data distributions. In doing so, we also proposed a novel Polynomial Divergence in the loss function to replace or generalize the Kullback-Leibler (KL) divergence, which results in systematic and drastic improvements in both accuracy and reproducibility of the re-constructed distribution function as well as the quality of re-constructed data images while keeping the dataset size the same but capturing fine resolution of the data. Moreover, we showed that the proposed PH-VAE has some form of disentangled representation learning ability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。