用几何视角重解变分自编码器,揭示隐空间如何通过正则化实现高效生成
From Points to Spheres: A Geometric Reinterpretation of Variational Autoencoders
- 将隐变量视为高斯球而非点,利用KL散度约束构建语义流形
- 重参数化机制在编解码间建立关键协作关系,支持从随机区域重建
- 统一解释VQ-VAE为聚类中心约束的生成模型,强调紧凑性而非随机性
变分自编码器通常从概率推断角度理解。本文提出一种新的几何诠释,补充概率视角并增强直观性。我们证明,语义流形的合理构造主要源于编码器中KL散度的约束作用。将潜在表示视为高斯球而非确定点,在KL散度约束下,高斯球对隐空间进行正则化,促进编码分布更均匀。此外,我们表明重参数化在编码器与解码器之间建立了关键的契约机制,使解码器能够从这些随机区域学习重构。我们进一步将该视角与VQ-VAE联系起来,提供统一理解:VQ-VAE可被视为将编码约束于一组聚类中心的自编码器,其生成能力源于紧凑性而非随机性。这一几何框架为理解VAE如何塑造隐空间几何以实现有效生成提供了新视角。
原文摘要 · Abstract (English)
Variational Autoencoder is typically understood from the perspective of probabilistic inference. In this work, we propose a new geometric reinterpretation which complements the probabilistic view and enhances its intuitiveness. We demonstrate that the proper construction of semantic manifolds arises primarily from the constraining effect of the KL divergence on the encoder. We view the latent representations as a Gaussian ball rather than deterministic points. Under the constraint of KL divergence, Gaussian ball regularizes the latent space, promoting a more uniform distribution of encodings. Furthermore, we show that reparameterization establishes a critical contractual mechanism between the encoder and decoder, enabling the decoder to learn how to reconstruct from these stochastic regions. We further connect this viewpoint with VQ-VAE, offering a unified perspective: VQ-VAE can be seen as an autoencoder where encodings are constrained to a set of cluster centers, with its generative capability arising from the compactness rather than its stochasticity. This geometric framework provides a new lens for understanding how VAE shapes the latent geometry to enable effective generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。