arXiv:2412.14865cs.LG2024-12被引 3

提出分层策略子空间,让智能体持续学习新导航任务时不忘旧技能。

Hierarchical Subspaces of Policies for Continual Offline Reinforcement Learning

  • 用神经网络的分层策略子空间实现灵活高效的任务适应。
  • 在MuJoCo和游戏类导航环境中表现优异,内存占用低且效率高。
  • 适合需要长期积累技能的机器人与游戏智能体场景。

我们研究持续强化学习场景,其中学习智能体需持续适应新任务,同时保留已有技能,重点解决遗忘旧知识及任务数量增长带来的可扩展性问题。这类挑战在自主机器人与视频游戏模拟中尤为突出,尤其在易受拓扑或运动学变化影响的导航任务中。为此,我们提出HiSPO,一种专为离线数据下导航任务设计的新型分层框架。该方法利用神经网络中不同的策略子空间,实现对新任务的灵活高效适应,同时保留已有知识。通过系统的实验验证,我们的方法在经典MuJoCo迷宫环境和复杂视频游戏类导航模拟中均表现出色,具备竞争力的性能和良好的适应性,在传统持续学习指标上尤其体现在内存使用少、效率高。

原文摘要 · Abstract (English)

We consider a Continual Reinforcement Learning setup, where a learning agent must continuously adapt to new tasks while retaining previously acquired skill sets, with a focus on the challenge of avoiding forgetting past gathered knowledge and ensuring scalability with the growing number of tasks. Such issues prevail in autonomous robotics and video game simulations, notably for navigation tasks prone to topological or kinematic changes. To address these issues, we introduce HiSPO, a novel hierarchical framework designed specifically for continual learning in navigation settings from offline data. Our method leverages distinct policy subspaces of neural networks to enable flexible and efficient adaptation to new tasks while preserving existing knowledge. We demonstrate, through a careful experimental study, the effectiveness of our method in both classical MuJoCo maze environments and complex video game-like navigation simulations, showcasing competitive performances and satisfying adaptability with respect to classical continual learning metrics, in particular regarding the memory usage and efficiency.

持续学习策略子空间离线强化学习导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。