arXiv:2409.14851cs.CVcs.LG2024-09中稿 · Neurocomputing被引 4

用离散编码提升表征解耦效果,无需真实因子信息

Disentanglement with Factor Quantized Variational Autoencoders

  • 通过标量量化构建全局离散潜变量,实现因子解耦
  • 在DCI和InfoMEC上优于现有方法,重建性能更优
  • 适合对表征可解释性要求高的生成模型研究者

解耦表征学习旨在将数据的潜在生成因子独立地表示在潜空间中。本文提出一种基于离散变分自编码器(VAE)的模型,无需提供生成因子的真实信息。我们证明了学习离散表示相比连续表示更能促进解耦。进一步地,我们在模型中引入归纳偏置:将潜变量进行标量量化,使用全局码本中的离散值,并在优化目标中加入总相关性项作为归纳偏置。所提出的FactorQVAE方法结合了基于优化的解耦与离散表示学习,在两个解耦度量(DCI 和 InfoMEC)上均优于现有方法,同时提升了重建性能。代码已开源:https://github.com/ituvisionlab/FactorQVAE。

原文摘要 · Abstract (English)

Disentangled representation learning aims to represent the underlying generative factors of a dataset in a latent representation independently of one another. In our work, we propose a discrete variational autoencoder (VAE) based model where the ground truth information about the generative factors are not provided to the model. We demonstrate the advantages of learning discrete representations over learning continuous representations in facilitating disentanglement. Furthermore, we propose incorporating an inductive bias into the model to further enhance disentanglement. Precisely, we propose scalar quantization of the latent variables in a latent representation with scalar values from a global codebook, and we add a total correlation term to the optimization as an inductive bias. Our method called FactorQVAE combines optimization based disentanglement approaches with discrete representation learning, and it outperforms the former disentanglement methods in terms of two disentanglement metrics (DCI and InfoMEC) while improving the reconstruction performance. Our code can be found at https://github.com/ituvisionlab/FactorQVAE.

解耦表征离散编码变分自编码器生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。