通过隐式熵优化解决变分自编码器后验崩溃问题
Entropic Auto-Encoding via Implicit Free-Energy Minimization

- 仅用重构损失,靠熵驱动隐变量先验生成
- 学习非高斯多模态隐空间,生成多样且数据一致
- 适合需结构化表示与隐含类别发现的任务
尽管变分自编码器(VAEs)广泛应用,但其固有的后验崩溃问题导致隐变量被忽略。该问题源于显式先验施加使优化趋向无信息隐表示的损失景观区域。本文提出熵自动编码器(EAE),其中重建损失是唯一显式目标,熵通过解码器集合的自由能最小化机制隐式生成隐变量先验。该集合使学习偏向近优解的高体积区域,而解码器更新引导搜索轨迹向有信息的隐表示。实验表明,EAE通过学习非高斯、多模态隐分布,实现多样化且数据一致的生成,并保留数据中的不同结构。作为概念验证,EAE捕捉了反应-扩散过程已知低维动态的叠加态;在MNIST上识别出隐含类别差异;在CelebA上展现出从“全人类面孔”到个体特征的层次化理解。
原文摘要 · Abstract (English)
Despite their ubiquity, variational autoencoders (VAEs) inherently suffer from posterior collapse, a failure mode in which latent variables are effectively ignored. This failure arises because explicit prior imposition drives optimization toward loss landscape regions corresponding to uninformative latent representations. Here, we introduce Entropic Autoencoders (EAEs), a framework in which reconstruction loss is the only explicit objective, and entropy generates the latent variables' prior implicitly through a free energy-minimizing ensemble of encoders. This ensemble biases learning toward high-volume regions of near-optimal solutions, while decoder updates direct the search trajectories toward informative latent representations. We demonstrate that EAEs mitigate posterior collapse by learning non-Gaussian, multimodal latent distributions that yield diverse, data-consistent generations and preserve different forms of underlying structure in the data. As a proof-of-concept, we show that an EAE captures a superposition of the known low-dimensional dynamics of a reaction-diffusion process. Then, we show that an EAE identifies implicit categorical distinctions in MNIST latent representations, and displays a hierarchical understanding of facial structure on the CelebA dataset, from an "all-human" face to individual-dependent features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。