通过动态正交性保持网络持续学习时的可塑性
Preserving Plasticity in Continual Learning via Dynamical Isometry

- 用动态正交性约束层间雅可比奇异值接近1,维持学习能力
- 新优化器AdamO可激活沉睡的ReLU单元,提升模型适应性
- 适用于需长期学习的场景,如机器人控制与在线训练
在非平稳环境下持续训练深度神经网络常导致可塑性逐渐丧失,阻碍进一步学习。本文将可塑性与经验神经正切核关联,发现动态正交性(即各层雅可比奇异值接近1)是维持持续学习可塑性的关键机制。研究重新审视一类几乎处处正交且仍具通用利普希茨逼近能力的网络,证明近似动态正交性与强非线性表达能力兼容。针对一般架构,提出高效正交性促进正则化方案,并揭示其通过重激活沉睡ReLU单元的新机制。基于此,提出类Adam自适应优化器AdamO,将正交性正则与梯度更新解耦,类似AdamW。进一步从动态正交性视角重新解释已有可塑性保护方法,发现它们仅作用于部分正交性指标。在设计用于诱发可塑性下降的监督与强化学习持续学习基准上,本方法表现稳定优于或匹配现有方法。
原文摘要 · Abstract (English)
Continual training of deep neural networks under non-stationarity often leads to a progressive loss of plasticity, eventually limiting further learning. We relate plasticity to the empirical Neural Tangent Kernel, and identify dynamical isometry (the condition that layer-wise Jacobian singular values remain close to one) as a key mechanism for preserving plasticity in continual learning. We revisit a class of networks that are almost-everywhere isometric while remaining universal Lipschitz function approximators, demonstrating that near-dynamical isometry is compatible with expressive nonlinear representations. For general architectures, we propose an efficient isometry-promoting regularization scheme and identify a novel mechanism by which it can reactivate dormant ReLU units. Building on this, we introduce AdamO, an Adam-style adaptive optimizer that decouples isometry regularization from gradient updates, analogous to AdamW. We further reinterpret prior plasticity-preserving approaches through the lens of dynamical isometry, showing that they target only a partial measure of isometry. Across supervised and reinforcement-learning continual-learning benchmarks designed to induce plasticity loss, our methods consistently match or outperform existing approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。