arXiv:2509.05545cs.LG2025-09被引 3

提出一种分层强化学习框架,让智能体提前规划路径,更稳定地完成长程任务。

Reinforcement Learning with Anticipation: A Hierarchical Approach for Long-Horizon Tasks

  • 分层设计:低层策略达成子目标,高层预测最优中间目标。
  • 理论保证:在多种条件下逼近全局最优策略,训练更稳定。
  • 适合研究者:关注长时序任务、分层决策与可证明性方法的学者。

解决长时序目标条件化任务仍是强化学习中的重大挑战。分层强化学习通过将任务分解为更易处理的子任务来应对,但自动发现层次结构以及多级策略联合训练常因不稳定性且缺乏理论保障而受限。本文提出强化学习中的预见机制(RLA),一种原则性强且可能可扩展的框架。RLA智能体学习两个协同模型:低层目标条件策略,用于到达指定子目标;高层预见模型,作为规划器,提出通往最终目标的最优路径上的中间子目标。RLA的关键在于预见模型的训练,其基于价值几何一致性原则进行优化,并通过正则化防止退化解。我们证明了在各种条件下,RLA可逼近全局最优策略,为长时序目标条件化任务中的分层规划与执行提供了原则性且收敛的方法。

原文摘要 · Abstract (English)

Solving long-horizon goal-conditioned tasks remains a significant challenge in reinforcement learning (RL). Hierarchical reinforcement learning (HRL) addresses this by decomposing tasks into more manageable sub-tasks, but the automatic discovery of the hierarchy and the joint training of multi-level policies often suffer from instability and can lack theoretical guarantees. In this paper, we introduce Reinforcement Learning with Anticipation (RLA), a principled and potentially scalable framework designed to address these limitations. The RLA agent learns two synergistic models: a low-level, goal-conditioned policy that learns to reach specified subgoals, and a high-level anticipation model that functions as a planner, proposing intermediate subgoals on the optimal path to a final goal. The key feature of RLA is the training of the anticipation model, which is guided by a principle of value geometric consistency, regularized to prevent degenerate solutions. We present proofs that RLA approaches the globally optimal policy under various conditions, establishing a principled and convergent method for hierarchical planning and execution in long-horizon goal-conditioned tasks.

强化学习分层决策长时序任务规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。