通过自动发现通用动作,实现多任务快速适应的分层强化学习新架构。
Hierarchical Meta-Reinforcement Learning via Automated Macro-Action Discovery
- 分三层结构:任务表征、自动发现通用宏观动作、基础动作学习。
- 在MetaWorld上样本效率和成功率显著优于现有最优方法。
- 通用宏观动作可跨任务复用,适合复杂高维任务快速迁移场景。
元强化学习(Meta-RL)能够快速适应新测试任务。尽管已有进展,但在多个复杂且高维的任务上学习高效策略仍具挑战。为此,我们提出一种新型三层次架构:1)学习任务表征,2)自动化发现与任务无关的宏观动作,3)学习基础动作。宏观动作能引导底层基础动作策略更高效地到达目标状态,缓解学习新冲突任务时遗忘旧行为的问题。通过从状态空间中移除任务特定成分,宏观动作具备任务无关性,支持跨任务重组,从而实现优异的快速适应能力。此外,通过创新的独立训练方案有效缓解了三层架构可能带来的不稳定性。在MetaWorld框架上的实验表明,该方法相比先前最先进方法在样本效率和成功率上均有提升。
原文摘要 · Abstract (English)
Meta-Reinforcement Learning (Meta-RL) enables fast adaptation to new testing tasks. Despite recent advancements, it is still challenging to learn performant policies across multiple complex and high-dimensional tasks. To address this, we propose a novel architecture with three hierarchical levels for 1) learning task representations, 2) discovering task-agnostic macro-actions in an automated manner, and 3) learning primitive actions. The macro-action can guide the low-level primitive policy learning to more efficiently transition to goal states. This can address the issue that the policy may forget previously learned behavior while learning new, conflicting tasks. Moreover, the task-agnostic nature of the macro-actions is enabled by removing task-specific components from the state space. Hence, this makes them amenable to re-composition across different tasks and leads to promising fast adaptation to new tasks. Also, the prospective instability from the tri-level hierarchies is effectively mitigated by our innovative, independently tailored training schemes. Experiments in the MetaWorld framework demonstrate the improved sample efficiency and success rate of our approach compared to previous state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。