arXiv:2410.09505cs.LGcs.NE2024-10被引 2

受海马体启发,提升长程规划的采样效率与泛化能力。

HG2P: Hippocampus-inspired High-reward Graph and Model-Free Q-Gradient Penalty for Path Planning and Motion Control

  • 基于高回报采样构建记忆图,提升样本效率
  • 提出无模型梯度惩罚,增强策略泛化性
  • 适用于复杂导航与机器人操作任务

目标条件下的分层强化学习(HRL)将复杂的抓取任务分解为一系列简单的子目标条件任务,在大规模环境中的长时程规划中展现出巨大潜力。本文借鉴脑机制,提出类海马-纹状体双控制器假说,结合生物体高回报偏好和实例理论,设计了一种高回报采样策略以构建记忆图,提升了样本效率。此外,推导出一种无模型的低层Q函数梯度惩罚,解决了先前方法中存在的模型依赖问题,改善了利普希茨约束在实际应用中的泛化能力。最终,将高回报图与无模型梯度惩罚(HG2P)整合进先进框架ACLG,提出新型目标条件分层强化学习框架HG2P+ACLG。实验表明,该方法在多种长时程导航和机器人操作任务中均优于当前最先进的目标条件HRL算法。

原文摘要 · Abstract (English)

Goal-conditioned hierarchical reinforcement learning (HRL) decomposes complex reaching tasks into a sequence of simple subgoal-conditioned tasks, showing significant promise for addressing long-horizon planning in large-scale environments. This paper bridges the goal-conditioned HRL based on graph-based planning to brain mechanisms, proposing a hippocampus-striatum-like dual-controller hypothesis. Inspired by the brain mechanisms of organisms (i.e., the high-reward preferences observed in hippocampal replay) and instance-based theory, we propose a high-return sampling strategy for constructing memory graphs, improving sample efficiency. Additionally, we derive a model-free lower-level Q-function gradient penalty to resolve the model dependency issues present in prior work, improving the generalization of Lipschitz constraints in applications. Finally, we integrate these two extensions, High-reward Graph and model-free Gradient Penalty (HG2P), into the state-of-the-art framework ACLG, proposing a novel goal-conditioned HRL framework, HG2P+ACLG. Experimentally, the results demonstrate that our method outperforms state-of-the-art goal-conditioned HRL algorithms on a variety of long-horizon navigation tasks and robotic manipulation tasks.

强化学习分层控制路径规划神经科学启发

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。