arXiv:2512.04958cs.LGcs.AI2025-12

提出可实现抽象框架,让分层强化学习获得近优解保证。

Realizable Abstractions: Near-Optimal Hierarchical Reinforcement Learning

  • 定义可实现抽象关系,避免非马尔可夫问题。
  • 任意抽象策略可转为低层近优策略,且通过约束MDP求解选项。
  • 算法RARL在多项式样本内收敛,对抽象误差鲁棒。

分层强化学习(HRL)旨在通过组合小任务的局部解,更高效地求解大型马尔可夫决策过程(MDP)。尽管直觉上合理,现有大多数MDP抽象概念表达能力有限或缺乏形式化效率保障。本文提出「可实现抽象」,建立通用低层MDP与其高层决策过程之间的新关系,避免非马尔可夫性问题,并具备理想的近优性保证。我们证明:任何针对该抽象的策略均可通过适当选项组合,转化为低层MDP的近优策略。这些选项可通过特定约束MDP求解。基于此,提出RARL算法,能输出可组合且近优的低层策略,利用输入的可实现抽象。理论证明:RARL是大概率近似正确(PAC),在多项式样本内收敛,且对抽象不准确性具有鲁棒性。

原文摘要 · Abstract (English)

The main focus of Hierarchical Reinforcement Learning (HRL) is studying how large Markov Decision Processes (MDPs) can be more efficiently solved when addressed in a modular way, by combining partial solutions computed for smaller subtasks. Despite their very intuitive role for learning, most notions of MDP abstractions proposed in the HRL literature have limited expressive power or do not possess formal efficiency guarantees. This work addresses these fundamental issues by defining Realizable Abstractions, a new relation between generic low-level MDPs and their associated high-level decision processes. The notion we propose avoids non-Markovianity issues and has desirable near-optimality guarantees. Indeed, we show that any abstract policy for Realizable Abstractions can be translated into near-optimal policies for the low-level MDP, through a suitable composition of options. As demonstrated in the paper, these options can be expressed as solutions of specific constrained MDPs. Based on these findings, we propose RARL, a new HRL algorithm that returns compositional and near-optimal low-level policies, taking advantage of the Realizable Abstraction given in the input. We show that RARL is Probably Approximately Correct, it converges in a polynomial number of samples, and it is robust to inaccuracies in the abstraction.

分层强化学习抽象近优性算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。