提出神经元级稳定与可塑性平衡机制,提升强化学习持续学习能力
Neuron-level Balance between Stability and Plasticity in Deep Reinforcement Learning
- 通过目标导向方法识别关键技能神经元,实现细粒度控制
- 在Meta-World和Atari上显著优于现有方法,兼顾旧技能保留与新任务适应
- 适合需要长期持续学习的强化学习场景
与人类持续获取知识的能力不同,深度强化学习(DRL)代理在稳定性与可塑性之间存在权衡:既要保留已有技能(稳定性),又要学习新知识(可塑性)。当前方法多在网络层面平衡二者,缺乏对单个神经元的区分与精细调控。为此,我们提出神经元级稳定与可塑性平衡(NBSP)方法,受特定神经元与任务相关技能高度关联的启发。具体而言,NBSP首先通过目标导向方法定义并识别强化学习中的关键技能神经元,用于知识保留;随后构建一个基于梯度掩码和经验回放的技术框架,针对这些神经元进行保护,以维持已有技能,同时支持对新任务的适应。在Meta-World和Atari基准上的大量实验表明,NBSP在平衡稳定性与可塑性方面显著优于现有方法。
原文摘要 · Abstract (English)
In contrast to the human ability to continuously acquire knowledge, agents struggle with the stability-plasticity dilemma in deep reinforcement learning (DRL), which refers to the trade-off between retaining existing skills (stability) and learning new knowledge (plasticity). Current methods focus on balancing these two aspects at the network level, lacking sufficient differentiation and fine-grained control of individual neurons. To overcome this limitation, we propose Neuron-level Balance between Stability and Plasticity (NBSP) method, by taking inspiration from the observation that specific neurons are strongly relevant to task-relevant skills. Specifically, NBSP first (1) defines and identifies RL skill neurons that are crucial for knowledge retention through a goal-oriented method, and then (2) introduces a framework by employing gradient masking and experience replay techniques targeting these neurons to preserve the encoded existing skills while enabling adaptation to new tasks. Numerous experimental results on the Meta-World and Atari benchmarks demonstrate that NBSP significantly outperforms existing approaches in balancing stability and plasticity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。