自动发明可复用的高层策略,提升持续强化学习的泛化与效率
Autonomous Option Invention for Continual Hierarchical Reinforcement Learning and Planning
- 通过可解释的状态抽象,自动生成符号化高层选项
- 在长周期稀疏奖励任务中实现比现有方法更优的样本效率
- 适合需要持续学习与跨任务迁移的智能体系统
抽象是提升强化学习可扩展性的关键。然而,自主学习抽象状态与动作表示以实现迁移和泛化仍是开放难题。本文提出一种新颖方法,用于在持续强化学习场景中自动发明、表示和利用选项(即长时间行为片段)。该方法应对具有长时域、稀疏奖励及未知转移与奖励函数的随机问题流。通过持续学习并维护可解释的状态抽象,生成具有符号化表达的高层选项。这些选项满足三大要求:(1) 可组合性,支持前瞻规划高效解题;(2) 可复用性,在不同问题实例间减少重学需求;(3) 相互独立性,降低选项间的干扰。主要贡献包括持续学习可迁移、可泛化的符号化选项方法,以及将搜索技术与强化学习结合,高效规划所学选项以解决新问题。实验表明,该方法能有效跨实例学习并传递抽象知识,相比最先进方法显著提升样本效率。
原文摘要 · Abstract (English)
Abstraction is key to scaling up reinforcement learning (RL). However, autonomously learning abstract state and action representations to enable transfer and generalization remains a challenging open problem. This paper presents a novel approach for inventing, representing, and utilizing options, which represent temporally extended behaviors, in continual RL settings. Our approach addresses streams of stochastic problems characterized by long horizons, sparse rewards, and unknown transition and reward functions. Our approach continually learns and maintains an interpretable state abstraction, and uses it to invent high-level options with abstract symbolic representations. These options meet three key desiderata: (1) composability for solving tasks effectively with lookahead planning, (2) reusability across problem instances for minimizing the need for relearning, and (3) mutual independence for reducing interference among options. Our main contributions are approaches for continually learning transferable, generalizable options with symbolic representations, and for integrating search techniques with RL to efficiently plan over these learned options to solve new problems. Empirical results demonstrate that the resulting approach effectively learns and transfers abstract knowledge across problem instances, achieving superior sample efficiency compared to state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。