arXiv:2411.01978cs.LGcs.AI2024-11中稿 · the Unifying Repre…被引 2

用内在维数和信息失衡分析VAE隐藏表征,揭示瓶颈大小的关键作用。

Understanding Variational Autoencoders with Intrinsic Dimension and Information Imbalance

  • 通过内在维数与信息失衡分析VAE隐藏层特征
  • 瓶颈超过数据内在维数时出现双峰维数曲线与信息处理突变
  • 可辅助模型架构搜索与过拟合诊断,适合生成模型研究者

本文利用内在维数(ID)和信息失衡(II)分析变分自编码器(VAE)的隐藏表示。结果表明,当瓶颈尺寸超过数据内在维数时,VAE行为发生转变,表现为双峰型内在维数分布,并在信息处理上出现定性差异。对于足够大的瓶颈,训练过程呈现两个阶段:快速拟合与缓慢泛化,由ID、II和KL散度的不同表现所体现。这些发现表明,ID与II可作为架构搜索的工具,用于诊断VAE欠拟合问题,并推动通过几何分析统一理解深度生成模型。

原文摘要 · Abstract (English)

This work presents an analysis of the hidden representations of Variational Autoencoders (VAEs) using the Intrinsic Dimension (ID) and the Information Imbalance (II). We show that VAEs undergo a transition in behaviour once the bottleneck size is larger than the ID of the data, manifesting in a double hunchback ID profile and a qualitative shift in information processing as captured by the II. Our results also highlight two distinct training phases for architectures with sufficiently large bottleneck sizes, consisting of a rapid fit and a slower generalisation, as assessed by a differentiated behaviour of ID, II, and KL loss. These insights demonstrate that II and ID could be valuable tools for aiding architecture search, for diagnosing underfitting in VAEs, and, more broadly, they contribute to advancing a unified understanding of deep generative models through geometric analysis.

变分自编码器内在维数信息失衡生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。