用因果图发现子目标结构,智能干预提升长时序强化学习效率
Hierarchical Reinforcement Learning with Targeted Causal Interventions
- 将子目标建模为因果图,自动发现层级结构
- 针对性干预重要子目标,训练成本显著降低
- 理论分析+实验验证,适用于树状与随机图结构
层次化强化学习(HRL)通过将任务分解为子目标层级来提升长时序稀疏奖励任务的效率。其核心挑战在于高效发现子目标间的层级结构并利用该结构达成最终目标。本文将子目标结构建模为因果图,并提出一种因果发现算法来学习该结构。不同于以往随机探索干预子目标的方式,我们基于已发现的因果模型,优先对在达成最终目标中起关键作用的子目标进行干预。这种有针对性的干预策略显著降低了训练成本。与先前缺乏理论分析的因果HRL工作不同,本文提供了形式化分析:对于树状结构及一类Erdős-Rényi随机图变体,本方法均表现优异。实验结果表明,该框架在多个HRL任务中优于现有方法,训练成本更低。
原文摘要 · Abstract (English)
Hierarchical reinforcement learning (HRL) improves the efficiency of long-horizon reinforcement-learning tasks with sparse rewards by decomposing the task into a hierarchy of subgoals. The main challenge of HRL is efficient discovery of the hierarchical structure among subgoals and utilizing this structure to achieve the final goal. We address this challenge by modeling the subgoal structure as a causal graph and propose a causal discovery algorithm to learn it. Additionally, rather than intervening on the subgoals at random during exploration, we harness the discovered causal model to prioritize subgoal interventions based on their importance in attaining the final goal. These targeted interventions result in a significantly more efficient policy in terms of the training cost. Unlike previous work on causal HRL, which lacked theoretical analysis, we provide a formal analysis of the problem. Specifically, for tree structures and, for a variant of Erdős-Rényi random graphs, our approach results in remarkable improvements. Our experimental results on HRL tasks also illustrate that our proposed framework outperforms existing work in terms of training cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。