arXiv:2603.21173cs.LGcs.AI2026-03

揭示深度强化学习中可塑性下降的根本原因,提出优化视角新解释。

Rethinking Plasticity in Deep Reinforcement Learning

  • 提出优化中心可塑性假说,指出旧任务最优解成新任务劣局部极小
  • 发现神经元休眠本质是梯度为零,且不同任务间可塑性差异显著
  • 解释参数约束为何能缓解可塑性损失,适合研究持续学习的学者

本文探究深度强化学习中可塑性丧失的根本机制,即神经网络在非平稳环境中丧失适应能力的问题。现有研究多依赖滞留神经元或有效秩等描述性指标,无法解释优化动态。我们提出优化中心可塑性(OCP)假说:旧任务的最优解在新任务中成为劣局部极小,导致参数被锁定,阻碍后续学习。理论证明神经元休眠等同于梯度为零状态,表明梯度缺失是休眠主因。实验显示可塑性损失高度任务特异;在某任务中休眠率高的网络,换到差异大的任务后性能可与随机初始化网络相当,说明网络容量未损,仅受优化景观抑制。此外,该假说解释了参数约束通过防止深度陷入局部极小而缓解可塑性损失的原因。在多种非平稳场景中验证,本研究为理解并恢复复杂强化学习场景中的网络可塑性提供了严谨的优化框架。

原文摘要 · Abstract (English)

This paper investigates the fundamental mechanisms driving plasticity loss in deep reinforcement learning (RL), a critical challenge where neural networks lose their ability to adapt to non-stationary environments. While existing research often relies on descriptive metrics like dormant neurons or effective rank, these summaries fail to explain the underlying optimization dynamics. We propose the Optimization-Centric Plasticity (OCP) hypothesis, which posits that plasticity loss arises because optimal points from previous tasks become poor local optima for new tasks, trapping parameters during task transitions and hindering subsequent learning. We theoretically establish the equivalence between neuron dormancy and zero-gradient states, demonstrating that the absence of gradient signals is the primary driver of dormancy. Our experiments reveal that plasticity loss is highly task-specific; notably, networks with high dormancy rates in one task can achieve performance parity with randomly initialized networks when switched to a significantly different task, suggesting that the network's capacity remains intact but is inhibited by the specific optimization landscape. Furthermore, our hypothesis elucidates why parameter constraints mitigate plasticity loss by preventing deep entrenchment in local optima. Validated across diverse non-stationary scenarios, our findings provide a rigorous optimization-based framework for understanding and restoring network plasticity in complex RL domains.

强化学习可塑性优化动态持续学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。