arXiv:2512.17091cs.LGcs.AI2025-12被引 1

融合强化学习与模型预测规划,实现高效决策。

Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making

  • 用强化学习动作指导模型预测采样,动态调整探索策略。
  • 在多个任务中提升成功率最高达72%,收敛速度加快2.1倍。
  • 适合复杂环境下的实时决策,如自动驾驶和机器人控制。

我们提出一种解决具有层次结构规划问题的新方法,将强化学习与模型预测控制(MPC)规划相结合。该方法巧妙地耦合两种规划范式:利用强化学习动作指导MPPI采样器,并自适应聚合MPPI样本以改进价值估计。该自适应过程在价值估计不确定性高时增强MPPI探索,提升训练鲁棒性与策略性能。实验在赛车驾驶、改良版Acrobot及带障碍物的Lunar Lander等任务中验证,结果表明该方法在数据效率和整体性能上均优于现有方法,任务成功率最高提升72%,收敛速度相比非自适应采样快2.1倍。

原文摘要 · Abstract (English)

We propose a new approach for solving planning problems with a hierarchical structure, fusing reinforcement learning and MPC planning. Our formulation tightly and elegantly couples the two planning paradigms. It leverages reinforcement learning actions to inform the MPPI sampler, and adaptively aggregates MPPI samples to inform the value estimation. The resulting adaptive process leverages further MPPI exploration where value estimates are uncertain, and improves training robustness and the overall resulting policies. This results in a robust planning approach that can handle complex planning problems and easily adapts to different applications, as demonstrated over several domains, including race driving, modified Acrobot, and Lunar Lander with added obstacles. Our results in these domains show better data efficiency and overall performance in terms of both rewards and task success, with up to a 72% increase in success rate compared to existing approaches, as well as accelerated convergence (x2.1) compared to non-adaptive sampling.

强化学习规划机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。