学习的不可逆性导致模型适应能力衰减,解释了持续学习失败的本质原因。
A Thermodynamic Theory of Learning Part II: Critical Period Closure and Continual Learning Failure
- 用参数分布传输建模学习过程,揭示其不可逆性对适应能力的限制。
- 发现任务间组合会降低可重构自由度,即使性能不变也渐失适应潜力。
- 提出容量阈值判据,指出遗忘源于动态重构能力耗尽而非缺乏多任务解。
有限时间内的学习本质上是不可逆的。在本系列第二部分中,我们表明不可逆性通过学习动力学的复合结构对未来的适应性施加几何限制。连续学习阶段以传输映射形式相乘,其雅可比矩阵构成半群,秩与奇异值具有次可乘性,因此动态可重配置的自由度只能减少。我们定义兼容有效秩(compatible effective rank)来量化模型仍可动态访问的任务保留方向的对数体积。尽管任务表现可能保持不变,但有限时间学习会逐步降低此重构能力。我们证明了持续学习的容量阈值准则:设 $ m_B $ 为新任务 B 在先前任务 A 的任务保留流形上的海森矩阵稳定秩,若 $ m_B $ 超过残余兼容有效秩,则任务 B 与任务 A 轨迹级不兼容;任何充分适应必然引发遗忘。因此,灾难性遗忘并非因缺乏多任务解,而是由组合学习动力学下重构能力的不可逆损耗所致。这确立了持续学习的轨迹级容量极限。
原文摘要 · Abstract (English)
Learning performed over finite time is inherently irreversible. In Part~I of this series, we modeled learning as a transport process in the space of parameter distributions and derived the Epistemic Speed Limit (ESL), which lower-bounds entropy production under finite-time dynamics. In this work (Part~II), we show that irreversibility imposes a geometric restriction on future adaptability through the compositional structure of learning dynamics. Successive learning phases compose multiplicatively as transport maps, and their Jacobians form a semigroup whose rank and singular values are submultiplicative. As a result, dynamically usable degrees of reconfiguration can only decrease under composition. We formalize the remaining adaptability of a model in terms of compatible effective rank, defined as the log-volume of task-preserving directions that remain dynamically accessible. Although task performance may remain unchanged, finite-time learning can progressively reduce this reconfiguration capacity. We prove a capacity-threshold criterion for continual learning: let m_B denote the stable rank of the Hessian of a new task B restricted to the task-preserving manifold of a previously learned task A. If m_B exceeds the residual compatible effective rank, then task B is trajectory-level incompatible with task A; any sufficient adaptation necessarily induces forgetting. Thus catastrophic forgetting arises not from the absence of multi-task solutions, but from irreversible loss of reconfiguration capacity under compositional learning dynamics. This establishes a trajectory-level capacity limit for continual learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。