通过高斯算子迹与核范数,统一优化自编码器与混合密度网络的密度分解。
Contrastive Entropy Bounds for Density and Conditional Density Decomposition
- 用高斯算子迹和核范数分别替代传统目标函数,实现特征可解释性提升。
- 在小方差高斯混合数据下,上界可量化分析,样本多样性显著增强。
- 适合研究神经网络表征、密度建模及生成模型的开发者与研究人员。
本文从贝叶斯高斯视角研究神经网络特征的可解释性,指出优化目标即达到概率上界;学习模型旨在逼近使上界紧致且代价最优的密度,常以高斯混合密度形式呈现。以边缘分布的混合密度网络(MDN)和条件分布的自编码器为例。已知仅对自编码器而言,最小化输入输出误差等价于最大化输入与隐层间的依赖性。针对多输出网络生成多个中心构成高斯混合的情形,本文利用希尔伯特空间与分解方法进行分析。首次发现:自编码器目标等价于最大化高斯算子的迹(即正交基下各特征值之和)。当一对一映射非必需时,可改用最大化该算子的核范数(即奇异值之和),以提升整体秩。因此,迹可用于训练自编码器,核范数可用作训练MDN的分歧度量。第二部分采用希尔伯特空间内积与范数定义上下界,相比基于KL的边界多一额外范数,虽增加采样开销但有效提升样本多样性,避免多输出网络输出恒定值的平凡解。提出编码器-混合-解码器架构,解码器为多输出,每样本生成多个中心,可能进一步收紧上界。假设数据为小方差高斯混合,该上界可被定量追踪与分析。
原文摘要 · Abstract (English)
This paper studies the interpretability of neural network features from a Bayesian Gaussian view, where optimizing a cost is reaching a probabilistic bound; learning a model approximates a density that makes the bound tight and the cost optimal, often with a Gaussian mixture density. The two examples are Mixture Density Networks (MDNs) using the bound for the marginal and autoencoders using the conditional bound. It is a known result, not only for autoencoders, that minimizing the error between inputs and outputs maximizes the dependence between inputs and the middle. We use Hilbert space and decomposition to address cases where a multiple-output network produces multiple centers defining a Gaussian mixture. Our first finding is that an autoencoder's objective is equivalent to maximizing the trace of a Gaussian operator, the sum of eigenvalues under bases orthonormal w.r.t. the data and model distributions. This suggests that, when a one-to-one correspondence as needed in autoencoders is unnecessary, we can instead maximize the nuclear norm of this operator, the sum of singular values, to maximize overall rank rather than trace. Thus the trace of a Gaussian operator can be used to train autoencoders, and its nuclear norm can be used as divergence to train MDNs. Our second test uses inner products and norms in a Hilbert space to define bounds and costs. Such bounds often have an extra norm compared to KL-based bounds, which increases sample diversity and prevents the trivial solution where a multiple-output network produces the same constant, at the cost of requiring a sample batch to estimate and optimize. We propose an encoder-mixture-decoder architecture whose decoder is multiple-output, producing multiple centers per sample, potentially tightening the bound. Assuming the data are small-variance Gaussian mixtures, this upper bound can be tracked and analyzed quantitatively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。