精心设计激活函数可有效防止持续学习中的适应力退化。
Activation Function Design Sustains Plasticity in Continual Learning
- 通过分析负区形状与饱和特性,设计两种即插即用的非线性函数。
- 在两类任务中验证:分类增量与非平稳强化学习环境,均提升适应能力。
- 无需额外参数或任务调优,适合追求轻量高效的持续学习应用。
在独立同分布(i.i.d.)训练中,激活函数差异常因模型规模与优化调整而缩小。但在持续学习中,模型除灾难性遗忘外,还会逐渐丧失适应新任务的能力(称作适应力损失),而这种失效模式中非线性的作用尚未被充分探索。本文表明,激活函数选择是缓解适应力损失的关键、与架构无关的调控手段。基于对负区形状和饱和行为的属性级分析,提出两种即插即用的非线性函数(Smooth-Leaky 和 Randomized Smooth-Leaky),并在两类互补设置下进行评估:(i) 监督式类别增量基准;(ii) 旨在诱发可控分布与动态变化的非平稳 MuJoCo 强化学习环境。同时提供一种简单压力测试协议与诊断工具,将激活函数形状与适应性能关联。核心结论明确:精心设计激活函数可在不增加容量或任务特定调优的前提下,以轻量方式维持持续学习中的适应力。
原文摘要 · Abstract (English)
In independent, identically distributed (i.i.d.) training regimes, activation functions have been benchmarked extensively, and their differences often shrink once model size and optimization are tuned. In continual learning, however, the picture is different: beyond catastrophic forgetting, models can progressively lose the ability to adapt (referred to as loss of plasticity) and the role of the non-linearity in this failure mode remains underexplored. We show that activation choice is a primary, architecture-agnostic lever for mitigating plasticity loss. Building on a property-level analysis of negative-branch shape and saturation behavior, we introduce two drop-in nonlinearities (Smooth-Leaky and Randomized Smooth-Leaky) and evaluate them in two complementary settings: (i) supervised class-incremental benchmarks and (ii) reinforcement learning with non-stationary MuJoCo environments designed to induce controlled distribution and dynamics shifts. We also provide a simple stress protocol and diagnostics that link the shape of the activation to the adaptation under change. The takeaway is straightforward: thoughtful activation design offers a lightweight, domain-general way to sustain plasticity in continual learning without extra capacity or task-specific tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。