提出OTA方法,提升离线目标导向强化学习的长时序任务表现
Option-aware Temporally Abstracted Value for Offline Goal-Conditioned Reinforcement Learning
- 引入选项感知的时间抽象价值学习,优化高层策略的价值函数
- 在OGBench上实现复杂迷宫与视觉机器人操作任务的高性能
- 特别适合解决长周期目标导向任务中的优势估计偏差问题
离线目标导向强化学习(GCRL)可从大量状态-动作轨迹数据中训练目标达成策略,无需额外环境交互。然而,即使采用分层策略结构(如HIQL),长时序任务仍存在性能瓶颈。我们发现两大根源:高层策略难以生成合适子目标;长时序下优势估计符号频繁错误。为此,本文提出简单有效的方案——选项感知的时间抽象价值学习(OTA),将时间抽象融入时序差分学习过程。通过使价值更新具备选项感知能力,压缩有效时域长度,从而在长时序场景下获得更清晰的优势估计。实验表明,使用OTA价值函数训练的高层策略在最近提出的离线GCRL基准OGBench上表现优异,涵盖迷宫导航和视觉机器人操作等复杂任务。
原文摘要 · Abstract (English)
Offline goal-conditioned reinforcement learning (GCRL) offers a practical learning paradigm in which goal-reaching policies are trained from abundant state-action trajectory datasets without additional environment interaction. However, offline GCRL still struggles with long-horizon tasks, even with recent advances that employ hierarchical policy structures, such as HIQL. Identifying the root cause of this challenge, we observe the following insight. Firstly, performance bottlenecks mainly stem from the high-level policy's inability to generate appropriate subgoals. Secondly, when learning the high-level policy in the long-horizon regime, the sign of the advantage estimate frequently becomes incorrect. Thus, we argue that improving the value function to produce a clear advantage estimate for learning the high-level policy is essential. In this paper, we propose a simple yet effective solution: Option-aware Temporally Abstracted value learning, dubbed OTA, which incorporates temporal abstraction into the temporal-difference learning process. By modifying the value update to be option-aware, our approach contracts the effective horizon length, enabling better advantage estimates even in long-horizon regimes. We experimentally show that the high-level policy learned using the OTA value function achieves strong performance on complex tasks from OGBench, a recently proposed offline GCRL benchmark, including maze navigation and visual robotic manipulation environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。