arXiv:2606.21593cs.LGcs.IT2026-06

揭示深度学习表征中信息压缩与几何压缩的复杂关系

Geometric and Information Compression of Representations in Deep Learning

论文配图:Geometric and Information Compression of Representations in Deep Learning
图 1 · 摘自论文原文
  • 用类别聚类衡量几何压缩,结合条件熵瓶颈网络分析信息压缩
  • 发现低互信息不必然导致几何压缩,二者关系呈非线性反转
  • 提出泛化能力可能是二者关联的潜在混淆因素,适合研究表征理论者

深度神经网络将输入数据转化为支持多种下游任务的隐含表征,这些表征可在信息论和几何维度上刻画,但两者关系尚不清楚。核心问题在于:输入与表征间的低互信息是否必然意味着几何压缩,反之亦然?本文通过类别聚类衡量几何压缩,并在条件熵瓶颈(CEB)网络和连续丢弃网络中使用理论可靠的互信息估计方法进行研究。在受控噪声注入条件下评估互信息、几何压缩与泛化性能之间的相互作用。结果表明,低互信息并不稳定对应几何压缩,二者关系比通常假设的更为复杂。实验揭示二者存在负向非线性关系,且在不同训练设置下可能反转。结果支持一个假说:泛化能力可能是该关联的潜在混淆因子,而非其直接后果。

原文摘要 · Abstract (English)

Deep neural networks transform input data into latent representations that support a wide range of downstream tasks. These representations can be characterized along information-theoretic and geometric dimensions, but their relationship remains poorly understood. A central open question is whether low mutual information (MI) between inputs and representations necessarily implies geometrically compressed latent spaces and vice versa. We investigate this question using class-wise clustering as a measure of geometric compression and theoretically sound MI estimation in conditional entropy bottleneck (CEB) networks and continuous dropout networks. We evaluate the interplay between MI, geometric compression, and generalization on classification tasks under controlled noise injection schemes. Our findings show that low MI does not reliably correspond to geometric compression, and that the connection between the two is more nuanced than often assumed. Indeed, our experiments reveal a negative and nonlinear relationship that can reverse when varying training setup. Our results put forward a hypothesis that generalization acts as a potential confounder in this connection rather than being their direct consequence.

表征压缩信息论几何结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。