用热力学与信息几何统一解释正则化,揭示最优学习的几何本质。
Thermodynamically Optimal Regularization under Information-Geometric Constraints
- 基于信息几何和准静态过程假设,推导出热力学最优正则化的统一框架。
- 最优正则化对应最小化信念空间中的Fisher-Rao距离,具唯一性。
- 为模型训练效率提供可验证预测,适合关注理论基础的研究者。
现代机器学习依赖一系列经验成功的正则化方法(如权重衰减、Dropout、指数移动平均),但其理论基础各异。随着大模型训练能耗攀升,学习算法是否接近根本效率极限成为关键问题。本文提出一个统一理论框架,连接热力学最优性、信息几何与正则化。在三个明确假设下——(A1) 最优性需参数无关的信息度量;(A2) 信念状态为已知约束下的最大熵分布;(A3) 最优过程为准静态——证明了条件最优定理:信念空间上唯一的可接受几何为Fisher-Rao度量,热力学最优正则化等价于最小化到参考状态的平方Fisher-Rao距离。推导出高斯与圆形信念模型的诱导几何,分别对应双曲与冯·米塞斯流形,并表明经典正则化无法保证热力学最优性。引入学习的热力学效率概念,并提出可实验验证的预测。本工作为机器学习正则化提供了原理性的几何与热力学基础。
原文摘要 · Abstract (English)
Modern machine learning relies on a collection of empirically successful but theoretically heterogeneous regularization techniques, such as weight decay, dropout, and exponential moving averages. At the same time, the rapidly increasing energetic cost of training large models raises the question of whether learning algorithms approach any fundamental efficiency bound. In this work, we propose a unifying theoretical framework connecting thermodynamic optimality, information geometry, and regularization. Under three explicit assumptions -- (A1) that optimality requires an intrinsic, parametrization-invariant measure of information, (A2) that belief states are modeled by maximum-entropy distributions under known constraints, and (A3) that optimal processes are quasi-static -- we prove a conditional optimality theorem. Specifically, the Fisher--Rao metric is the unique admissible geometry on belief space, and thermodynamically optimal regularization corresponds to minimizing squared Fisher--Rao distance to a reference state. We derive the induced geometries for Gaussian and circular belief models, yielding hyperbolic and von Mises manifolds, respectively, and show that classical regularization schemes are structurally incapable of guaranteeing thermodynamic optimality. We introduce a notion of thermodynamic efficiency of learning and propose experimentally testable predictions. This work provides a principled geometric and thermodynamic foundation for regularization in machine learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。