arXiv:2605.22711cs.LGcs.AI2026-05

通过抽象层次结构提升离线目标导向强化学习的泛化能力

Abstraction for Offline Goal-Conditioned Reinforcement Learning

论文配图:Abstraction for Offline Goal-Conditioned Reinforcement Learning
图 1 · 摘自论文原文
  • 引入相对化选项与分层表征,实现状态空间跨场景复用
  • 在多个离线目标导向任务中性能显著优于基线方法
  • 适合研究离线强化学习、具身智能与复杂任务规划的学者

现实世界中的目标导向强化学习(GCRL)中,马尔可夫决策过程(MDPs)常因对称性与状态-目标对间的共享结构而存在显著冗余。尽管层级策略已被用于通过时间抽象减少规划时长,本文进一步证明层级结构还能实现绝对抽象。通过引入相对化选项以及分层表示,我们展示了智能体如何在相似的状态空间上下文中重用经验。基于此框架,提出两种简单算法以学习相对化选项并从绝对参考系中抽象。实验表明,此类归纳偏置在离线GCRL中能显著提升性能。

原文摘要 · Abstract (English)

Markov Decision Processes (MDPs) often exhibit significant redundancy due to symmetries and shared structure across state-goal pairs in real-world Goal-Conditioned Reinforcement Learning (GCRL). While hierarchical policies have been motivated for horizon reduction via temporal abstraction in offline GCRL, we demonstrate that hierarchy also enables absolute abstraction. By introducing relativised options as well as distinct representations for different levels of the hierarchy, we demonstrate how an agent can reuse experience across similar contexts of the state-space. Based on this framework, we introduce two simple algorithms for learning relativised options and abstracting from the absolute frame of reference. Our experiments show that such inductive biases significantly improve performance in offline GCRL.

强化学习离线学习层次策略抽象建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。