提出神经ODE微调不遗忘的几何理论基础,解释为何能精确保持旧任务性能。
Geometric Foundations of Tuning without Forgetting in Neural ODEs
- 基于控制函数子空间的微分几何分析,证明其为有限余维巴纳赫流形。
- 揭示微调过程本质是沿切空间连续变形控制函数,实现映射不变性。
- 为顺序训练中避免遗忘提供严格数学依据,适合研究模型稳定性者阅读。
在先前工作中,我们提出了神经ODE顺序训练中的无遗忘微调(TwF)原则,即在保留已学习样本输出映射的前提下,参数更新限定于控制函数的子空间内,在一阶近似意义下保持映射不变。本文证明:在非奇异控制条件下,该参数子空间构成有限余维的巴纳赫子流形,并刻画了其切空间。这表明TwF等价于沿该巴纳赫子流形的切空间对控制函数进行延续/变形,从而在精确意义上实现映射保持(不遗忘),超越了一阶近似的限制,为无遗忘机制提供了严格的理论支撑。
原文摘要 · Abstract (English)
In our earlier work, we introduced the principle of Tuning without Forgetting (TwF) for sequential training of neural ODEs, where training samples are added iteratively and parameters are updated within the subspace of control functions that preserves the end-point mapping at previously learned samples on the manifold of output labels in the first-order approximation sense. In this letter, we prove that this parameter subspace forms a Banach submanifold of finite codimension under nonsingular controls, and we characterize its tangent space. This reveals that TwF corresponds to a continuation/deformation of the control function along the tangent space of this Banach submanifold, providing a theoretical foundation for its mapping-preserving (not forgetting) during the sequential training exactly, beyond first-order approximation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。