发现深度网络训练末期特征流形维度下降,与信息压缩相关。
Learning to Compress: Local Rank and Information Compression in Deep Neural Networks
- 用局部秩衡量特征流形维度,揭示其随训练降低
- 证明网络在训练后期压缩输入与中间层的互信息
- 连接特征表示与信息瓶颈理论,适合理解模型压缩者
深度神经网络在训练过程中倾向于低秩解,隐式学习低维特征表示。本文研究深度多层感知机(MLPs)如何编码这些特征流形,并将其行为与信息瓶颈(IB)理论联系起来。引入局部秩作为特征流形维度的度量,理论上和实证上均表明该秩在训练最后阶段下降。我们认为,降低表示秩的网络也会压缩输入与中间层之间的互信息。这项工作弥合了特征流形秩与信息压缩之间的鸿沟,为信息瓶颈与表征学习的相互作用提供了新见解。
原文摘要 · Abstract (English)
Deep neural networks tend to exhibit a bias toward low-rank solutions during training, implicitly learning low-dimensional feature representations. This paper investigates how deep multilayer perceptrons (MLPs) encode these feature manifolds and connects this behavior to the Information Bottleneck (IB) theory. We introduce the concept of local rank as a measure of feature manifold dimensionality and demonstrate, both theoretically and empirically, that this rank decreases during the final phase of training. We argue that networks that reduce the rank of their learned representations also compress mutual information between inputs and intermediate layers. This work bridges the gap between feature manifold rank and information compression, offering new insights into the interplay between information bottlenecks and representation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。