揭示训练算法的不可逆性及其对学习轨迹的普遍影响
Thermodynamic Irreversibility of Training Algorithms

- 构建统一框架分析训练过程的不可逆性
- 四类不可逆性度量在步长下等价
- 不可逆性导致熵产率最小的轨迹被偏好
人工智能系统的训练算法均引入远离平衡的动态过程,理解这些算法的不可逆性是揭示现代AI系统学习动态的基础。本文建立了一个通用框架,用于定义和分析训练算法的不可逆性。我们证明了四种刻画动力学过程不可逆性的方法——数值反向误差ϕ_{ m DE}、时间重标度修正ϕ_{ m TR}、微观时间反演不对称性ϕ_{ m TA}以及(正则化)随机热力学熵产生ϕ_{ m ST}——在步长η的主导阶上等价。不可逆性引发一种破坏时间反演对称性的涌现力,通常破坏非保距连续重参数化对称性,保留正交对称性,并导致对熵产生率最小的学习轨迹的普遍偏好。
原文摘要 · Abstract (English)
The training algorithms for AI systems all introduce far-from-equilibrium dynamical processes, and understanding the irreversibility of these algorithms is a fundamental step towards understanding the learning dynamics of modern AI systems. In this work, we establish a general framework for defining and analyzing the irreversibility of training algorithms. We show that four different ways to characterize the irreversibility of dynamical processes are equivalent to leading order in the step size $η$: numerical backward error $ϕ_{\rm DE}$, time-renormalized correction $ϕ_{\rm TR}$, microscopic time reversal asymmetry $ϕ_{\rm TA}$, and the (regularized) stochastic-thermodynamic entropy production $ϕ_{\rm ST}$. The irreversibility gives rise to a time-reversal-symmetry-breaking emergent force that generically breaks non-isometric continuous reparametrization symmetries, preserves orthogonal symmetries, and leads to a universal preference for those learning trajectories that minimize the entropy production rate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。