提出新算法提升离线目标导向强化学习的长期任务表现
Latent Representation Alignment for Offline Goal-Conditioned Reinforcement Learning

- 用潜在表示对齐增强价值函数泛化能力
- 在22个数据集上20个领先,长程任务表现突出
- 适合研究离线强化学习与复杂任务规划的学者
离线目标导向强化学习(GCRL)为从固定数据集中获取目标达成策略提供了实用框架。然而,在长时序任务中学习可靠的条件价值函数仍具挑战。本文识别出目标条件价值函数的错误泛化是根本瓶颈,并证明价值函数中适当的归纳偏置至关重要。基于此,我们提出潜空间对齐价值学习(LAVL),将基于潜在表示的价值泛化与分层规划统一整合。在OGBench上的大量实验表明,LAVL始终优于现有离线GCRL方法,在22个数据集中有20个达到最高性能。尤其在长时序任务和轨迹拼接数据集上,其表现显著优于先前方法,而后者在此类任务中性能大幅下降。代码已公开于https://github.com/oh-lab/LAVL.git。
原文摘要 · Abstract (English)
Offline goal-conditioned reinforcement learning (GCRL) provides a practical framework for obtaining goal-reaching policies from fixed datasets. However, learning a reliable goal-conditioned value function in long-horizon tasks remains challenging. In this paper, we identify erroneous generalization in goal-conditioned value functions as a fundamental bottleneck, and demonstrate that appropriate inductive bias in the value function is crucial for addressing the bottleneck. Building on these findings, we propose Latent-Aligned Value Learning (LAVL), an offline GCRL algorithm that integrates latent-representation-based value generalization with hierarchical planning in a unified framework. Extensive experiments on OGBench demonstrate that LAVL consistently outperforms existing offline GCRL methods, achieving the highest performance on 20 out of 22 datasets. Notably, LAVL exhibits strong performance in long-horizon tasks and trajectory stitching datasets, where prior methods suffer significant performance degradation. Our code is available at https://github.com/oh-lab/LAVL.git.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。