揭示深度学习在持续学习中丧失适应能力的根本原因
Barriers for Learning in an Evolving World: Mathematical Understanding of Loss of Plasticity
- 从动力系统角度定义参数空间中的稳定流形为学习陷阱
- 发现激活饱和与表征冗余是导致学习停滞的两大机制
- 解释为何简化模型反而加剧持续学习中的适应障碍
深度学习模型在静态数据上表现优异,但在非平稳环境中因学习适应力下降(即损失可塑性,LoP)而失效。本文基于动力系统理论,首次对梯度学习中的LoP进行第一性原理分析。通过识别参数空间中的稳定流形,正式定义了阻碍梯度轨迹继续学习的陷阱。分析揭示两类主要陷阱形成机制:由激活饱和导致的冻结单元,以及由表征冗余引发的克隆单元流形。该框架揭示了一个根本矛盾:在静态场景中促进泛化的特性,如低秩表示和简单性偏好,在持续学习中反而加剧了可塑性丧失。研究通过数值模拟验证理论,并探索架构设计或针对性扰动作为缓解策略。
原文摘要 · Abstract (English)
Deep learning models excel in stationary data but struggle in non-stationary environments due to a phenomenon known as loss of plasticity (LoP), the degradation of their ability to learn in the future. This work presents a first-principles investigation of LoP in gradient-based learning. Grounded in dynamical systems theory, we formally define LoP by identifying stable manifolds in the parameter space that trap gradient trajectories. Our analysis reveals two primary mechanisms that create these traps: frozen units from activation saturation and cloned-unit manifolds from representational redundancy. Our framework uncovers a fundamental tension: properties that promote generalization in static settings, such as low-rank representations and simplicity biases, directly contribute to LoP in continual learning scenarios. We validate our theoretical analysis with numerical simulations and explore architectural choices or targeted perturbations as potential mitigation strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。