用范畴论理解训练动态,发现模型泛化能力由路径同伦性决定
Categorical Invariants of Learning Dynamics
- 将训练视为参数到表征的保结构变换,用同伦类分类优化路径
- 同伦路径收敛模型泛化误差差小于0.5%,非同伦路径超3%
- 提供可预测泛化的持久同调工具,适合研究泛化机制的学者
神经网络训练通常被视为在损失曲面上的梯度下降。我们提出一个根本不同的视角:学习是参数空间(Param)与表征空间(Rep)之间的结构保持变换(函子 L)。该范畴框架揭示,产生相似测试性能的不同训练过程通常属于同一同伦类(连续形变族)。实验表明,通过同伦路径收敛的网络泛化性能差异小于0.5%,而非同伦路径则超过3%。理论提供了实用工具:持久同调识别出可预测泛化的稳定极小值(相关系数 R²=0.82),拉回构造形式化了迁移学习,二范畴结构解释了不同优化算法何时产生功能等价模型。这些范畴不变量既揭示了深度学习有效的理论原因,也为训练更鲁棒网络提供了具体算法原则。
原文摘要 · Abstract (English)
Neural network training is typically viewed as gradient descent on a loss surface. We propose a fundamentally different perspective: learning is a structure-preserving transformation (a functor L) between the space of network parameters (Param) and the space of learned representations (Rep). This categorical framework reveals that different training runs producing similar test performance often belong to the same homotopy class (continuous deformation family) of optimization paths. We show experimentally that networks converging via homotopic trajectories generalize within 0.5% accuracy of each other, while non-homotopic paths differ by over 3%. The theory provides practical tools: persistent homology identifies stable minima predictive of generalization (R^2 = 0.82 correlation), pullback constructions formalize transfer learning, and 2-categorical structures explain when different optimization algorithms yield functionally equivalent models. These categorical invariants offer both theoretical insight into why deep learning works and concrete algorithmic principles for training more robust networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。