提出广义信息瓶颈理论,用协同信息提升模型泛化能力
A Generalized Information Bottleneck Theory of Deep Learning
- 用特征协同作用重新定义信息瓶颈,解决原理论不可计算问题
- 实验显示在多种架构中均出现压缩阶段,包括传统IB失效的ReLU网络
- 结果更符合对抗鲁棒性认知,适合研究深度学习泛化机制者阅读
信息瓶颈(IB)原理为理解神经网络学习提供了有力的理论框架,但其实际应用受限于未解的理论模糊性及难以准确估计的问题。本文提出一种广义信息瓶颈(GIB)框架,通过协同作用——即仅通过联合处理特征才能获得的信息——重构原始IB原理。我们提供了理论和实证证据,证明协同函数相比非协同函数具有更优的泛化性能。基于此,我们利用每个特征与其余特征的平均交互信息(II)构建可计算的协同定义,重新表述了IB。我们证明,在完美估计条件下,原始IB目标函数被我们的GIB所上界,确保与现有IB理论兼容的同时克服其局限性。实验表明,GIB在多种架构(包括标准IB失效的含ReLU激活函数网络)中均表现出一致的压缩阶段,且在卷积网络和Transformer中展现出可解释的动力学特性,并更贴近对抗鲁棒性的认知。
原文摘要 · Abstract (English)
The Information Bottleneck (IB) principle offers a compelling theoretical framework to understand how neural networks (NNs) learn. However, its practical utility has been constrained by unresolved theoretical ambiguities and significant challenges in accurate estimation. In this paper, we present a \textit{Generalized Information Bottleneck (GIB)} framework that reformulates the original IB principle through the lens of synergy, i.e., the information obtainable only through joint processing of features. We provide theoretical and empirical evidence demonstrating that synergistic functions achieve superior generalization compared to their non-synergistic counterparts. Building on these foundations we re-formulate the IB using a computable definition of synergy based on the average interaction information (II) of each feature with those remaining. We demonstrate that the original IB objective is upper bounded by our GIB in the case of perfect estimation, ensuring compatibility with existing IB theory while addressing its limitations. Our experimental results demonstrate that GIB consistently exhibits compression phases across a wide range of architectures (including those with \textit{ReLU} activations where the standard IB fails), while yielding interpretable dynamics in both CNNs and Transformers and aligning more closely with our understanding of adversarial robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。