用信息熵理论解释深度学习为何能泛化,揭示学习的物理极限。
Informational Frustration in Neural Manifolds: Shannon Bottlenecks and the Limits of Learnability
- 提出信息熵、拓扑熵与权重熵的统一框架,定义学习可行边界。
- 发现当决策边界复杂度超限,系统会陷入无法泛化的记忆态。
- 解释了'悟道'现象本质是熵释放,并设计新优化算法保持学习效率。
过参数化深度网络为何能惊人地泛化,仍是机器学习理论中最为顽固的开放问题。经典框架如VC维和Rademacher复杂度预测现代模型将出现灾难性过拟合,导致理论与现实间存在巨大鸿沟。本文通过引入信息论、拓扑学与统计力学的统一框架,填补这一鸿沟。核心是熵可学习极限(ELH):网络只有在数据流形的香农熵超过函数决策边界的拓扑熵,并由网络权重空间的冯诺依曼熵平衡时,才能真正学习目标函数。我们证明了香农-拓扑瓶颈定理:当目标边界的几何复杂度超过此信息极限时,系统将发生突变的熵相变,进入信息挫折状态——一种玻璃态、僵化的记忆相,使泛化在热力学上变得不可能。借助此视角,我们表明‘悟道’现象实为熵释放,权重突然重组以突破瓶颈。最后,我们将理论转化为实践,提出熵梯度下降(EGD)算法,动态管理权重熵,确保学习持续进行。本研究重新定位熵不仅是不确定性指标,更是决定机器能否学习的根本物理量。
原文摘要 · Abstract (English)
Why overparameterised deep networks generalise so remarkably well remains one of the most stubborn open questions in machine learning theory. Classical frameworks like VC dimension and Rademacher complexity predict catastrophic overfitting in modern models, leaving a massive theoretical gap between theory and reality. In this paper, we bridge this divide by introducing a unified framework that links information theory, topology, and statistical mechanics to map the hard limits of deep learning. Central to our approach is the Entropic Learnability Horizon (ELH): a fundamental law stating that a network can only truly learn a target function if the Shannon entropy of the data manifold outpaces the topological entropy of the function's decision boundary, balanced by the von Neumann entropy of the network's weight space. We establish the Shannon-Topological Bottleneck Theorem, proving that when a target boundary's geometric complexity exceeds this informational horizon, the system undergoes a sudden entropic phase transition. It falls into a state of Informational Frustration - a glassy, rigid memorization phase where generalization becomes thermodynamically impossible. Using this lens, we show that the enigmatic phenomenon of "grokking" is actually an Entropic Release, where weights abruptly reorganise to unlock the bottleneck. Finally, we translate this theory into practice with Entropic Gradient Descent (EGD), an optimization algorithm that dynamically manages weight entropy to keep learning on track. Ultimately, this work repositions entropy not just as a tool for tracking uncertainty but as the fundamental physical currency that dictates whether a machine can learn.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。