MANGO通过多层抽象实现选项嵌套,提升长周期稀疏奖励任务的样本效率与可解释性。
Multi-layer Abstraction for Nested Generation of Options (MANGO) in Hierarchical Reinforcement Learning
- 分层抽象生成嵌套选项,模块化宏观动作提升策略复用
- 在程序生成网格环境中样本效率提升显著,泛化能力更强
- 适合安全关键场景,决策过程逐层透明易理解
本文提出MANGO(多层抽象嵌套选项生成框架),一种新型层次强化学习方法,用于应对长期稀疏奖励环境中的挑战。MANGO将复杂任务分解为多层抽象,每层定义抽象状态空间,并使用选项将轨迹模块化为宏观动作。这些选项在各层间嵌套,实现已学动作的高效复用,提升样本效率。框架引入层内策略引导抽象状态转移,以及融合任务特定组件(如奖励函数)的任务动作。在程序生成的网格环境中实验表明,相比标准RL方法,MANGO在样本效率和泛化能力上均有显著提升。同时,其多层决策结构增强了可解释性,特别适用于安全敏感及工业场景。未来工作将探索自动发现抽象与动作、连续或模糊环境适应性,以及更鲁棒的多层训练策略。
原文摘要 · Abstract (English)
This paper introduces MANGO (Multilayer Abstraction for Nested Generation of Options), a novel hierarchical reinforcement learning framework designed to address the challenges of long-term sparse reward environments. MANGO decomposes complex tasks into multiple layers of abstraction, where each layer defines an abstract state space and employs options to modularize trajectories into macro-actions. These options are nested across layers, allowing for efficient reuse of learned movements and improved sample efficiency. The framework introduces intra-layer policies that guide the agent's transitions within the abstract state space, and task actions that integrate task-specific components such as reward functions. Experiments conducted in procedurally-generated grid environments demonstrate substantial improvements in both sample efficiency and generalization capabilities compared to standard RL methods. MANGO also enhances interpretability by making the agent's decision-making process transparent across layers, which is particularly valuable in safety-critical and industrial applications. Future work will explore automated discovery of abstractions and abstract actions, adaptation to continuous or fuzzy environments, and more robust multi-layer training strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。