arXiv:2605.03636cs.LG2026-05

分析二值神经网络的信息平面,发现压缩与泛化无固定关联。

Information Plane Analysis of Binary Neural Networks

论文配图:Information Plane Analysis of Binary Neural Networks
图 1 · 摘自论文原文
  • 用二值网络解决信息熵估计难题,确保信息平面分析可靠
  • 375个实验显示训练后期普遍有压缩现象但不提升泛化
  • 压缩与泛化关系受任务、结构和正则化影响,无普适规律

信息平面(IP)分析通过输入、表征与目标间的互信息(MI)研究深度神经网络的训练动态。然而,高维确定性表征样本的互信息估计常导致统计无效。本文对二值神经网络(BNNs)进行IP分析,其激活为离散值,互信息有限。我们刻画了插件熵估计的有限样本行为,识别出在样本数 $N$ 和表征维度 $D$ 下可信赖的估计区间。超出此区间时,经验互信息会饱和至 $"log_2 N$,使信息平面轨迹失去意义。限定在可靠区间内,我们训练了375个BNN,探究晚期压缩阶段的存在性及其与泛化性能的关系。结果表明,尽管晚期压缩常被观察到,但压缩的潜在表征并不一致地提升泛化性能。压缩与泛化的关系高度依赖于任务、架构和正则化方式。

原文摘要 · Abstract (English)

Information plane (IP) analysis has been suggested to study the training dynamics of deep neural networks through mutual information (MI) between inputs, representations, and targets. However, its statistical validity is often compromised by the difficulty of estimating MI from samples of high-dimensional, deterministic representations. In this work, we perform IP analyses on binary neural networks (BNNs) where activations are discrete and MI is finite. We characterise the finite-sample behaviour of the plug-in entropy estimator and identify regimes for sample size $N$ and representation dimensionality $D$ under which MI estimates are reliable. Outside these regimes, we show that empirical MI estimates saturate to $\log_2 N$, rendering IP trajectories uninformative. Restricting attention to the reliable regime, we train 375 BNNs to investigate the existence of late-stage compression phases and the relationship between compressed representations and generalisation performance. Our results show that while late-stage compression is frequently observed, compressed latent representations do not consistently correlate with improved generalization performance. Instead, the relationship between compression and generalisation is highly dependent on task, architecture, and regularisation.

信息平面二值网络泛化分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。