arXiv:2411.03008cs.LGcs.AI2024-11中稿 · as a poster被引 2

HOP通过观察相似性动态构建策略层级,避免持续学习中的灾难性遗忘。

Hierarchical Orchestra of Policies

  • 基于观察相似性自动构建策略层级,无需任务标签
  • 在可程序生成环境中显著提升跨任务知识保留能力
  • 适合任务边界模糊的连续学习场景,性能稳定

持续强化学习面临的主要挑战是代理在学习顺序任务时易产生灾难性遗忘。本文提出一种基于模块化的解决方案——分层策略乐团(Hierarchical Orchestra of Policies, HOP),旨在缓解终身强化学习中的灾难性遗忘问题。HOP根据当前观察与以往成功任务中观察的相似性,动态构建策略层次结构。与现有先进方法不同,HOP无需任务标注,可在任务边界模糊的环境中实现稳健适应。实验在多个可程序生成环境的任务序列上进行,结果表明HOP显著优于基线方法,在知识保留方面表现优异,且与需任务标注的先进迁移方法性能相当。此外,当任务保持不变时,HOP亦未牺牲性能,展现出良好泛化性。

原文摘要 · Abstract (English)

Continual reinforcement learning poses a major challenge due to the tendency of agents to experience catastrophic forgetting when learning sequential tasks. In this paper, we introduce a modularity-based approach, called Hierarchical Orchestra of Policies (HOP), designed to mitigate catastrophic forgetting in lifelong reinforcement learning. HOP dynamically forms a hierarchy of policies based on a similarity metric between the current observations and previously encountered observations in successful tasks. Unlike other state-of-the-art methods, HOP does not require task labelling, allowing for robust adaptation in environments where boundaries between tasks are ambiguous. Our experiments, conducted across multiple tasks in a procedurally generated suite of environments, demonstrate that HOP significantly outperforms baseline methods in retaining knowledge across tasks and performs comparably to state-of-the-art transfer methods that require task labelling. Moreover, HOP achieves this without compromising performance when tasks remain constant, highlighting its versatility.

持续学习策略层级灾难性遗忘无监督任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。