arXiv:2410.07994cs.LG2024-10ICLR被引 20

让AI模型像大脑一样持续学习,动态扩展网络结构。

Neuroplastic Expansion in Deep Reinforcement Learning

  • 通过梯度引导拓扑生成,动态扩张网络规模。
  • 在MuJoCo和DeepMind Control中超越现有方法,提升适应性。
  • 适合需要持续学习的复杂动态环境任务。

学习智能体的可塑性丧失(类似生物大脑神经通路固化)会显著阻碍强化学习中的学习与适应能力,因其非平稳特性。为应对这一根本挑战,我们提出一种受认知科学中皮层扩张启发的新方法——神经可塑性扩展(Neuroplastic Expansion, NE)。NE通过从初始小规模网络动态增长至完整维度,全程保持模型的可学习性与适应性。该方法包含三个核心组件:(1) 基于潜在梯度的弹性拓扑生成;(2) 通过休眠神经元剪枝优化网络表达力;(3) 通过经验回顾实现神经元整合,平衡可塑性与稳定性。大量实验表明,NE能有效缓解可塑性丧失,在MuJoCo和DeepMind Control Suite多个任务上优于当前最优方法。该方法使深度强化学习在复杂动态环境中具备更强自适应能力,是推动其从静态一次性训练迈向持续适应模型的重要一步。

原文摘要 · Abstract (English)

The loss of plasticity in learning agents, analogous to the solidification of neural pathways in biological brains, significantly impedes learning and adaptation in reinforcement learning due to its non-stationary nature. To address this fundamental challenge, we propose a novel approach, {\it Neuroplastic Expansion} (NE), inspired by cortical expansion in cognitive science. NE maintains learnability and adaptability throughout the entire training process by dynamically growing the network from a smaller initial size to its full dimension. Our method is designed with three key components: (\textit{1}) elastic topology generation based on potential gradients, (\textit{2}) dormant neuron pruning to optimize network expressivity, and (\textit{3}) neuron consolidation via experience review to strike a balance in the plasticity-stability dilemma. Extensive experiments demonstrate that NE effectively mitigates plasticity loss and outperforms state-of-the-art methods across various tasks in MuJoCo and DeepMind Control Suite environments. NE enables more adaptive learning in complex, dynamic environments, which represents a crucial step towards transitioning deep reinforcement learning from static, one-time training paradigms to more flexible, continually adapting models.

强化学习持续学习神经可塑性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。