arXiv:2509.20478cs.LG2025-09NeurIPS被引 18

统一对比与时序距离,实现更优的离线目标导向强化学习

Offline Goal-conditioned Reinforcement Learning with Quasimetric Representations

  • 结合对比学习与拟度量空间结构,构建可优化的目标可达距离
  • 在次优数据和随机环境中仍能学习最优目标到达策略
  • 适合处理高维噪声环境下的复杂目标规划任务

目标导向强化学习(GCRL)常依赖学习到的状态表示来提取目标达成策略。现有两种代表性结构框架表现优异:(1) 对比表示,通过对比目标学习‘后续特征’,实现对未来结果的推理;(2) 时间距离,将表示空间中的拟度量距离与状态到目标的转移时间关联。本文提出一种统一框架,利用拟度量空间结构(三角不等式)并引入额外约束,学习可实现最优目标达成的后续表示。相比以往方法,本方法即使在次优数据和随机环境下,也能利用拟度量参数化学习最优目标到达距离。该方法兼具蒙特卡洛对比学习的稳定性与长程能力,以及拟度量网络的自由拼接优势。在现有离线GCRL基准测试中,新表示学习目标在拼接任务上优于对比学习方法,在噪声高、维度高的环境中也超越拟度量网络方法。

原文摘要 · Abstract (English)

Approaches for goal-conditioned reinforcement learning (GCRL) often use learned state representations to extract goal-reaching policies. Two frameworks for representation structure have yielded particularly effective GCRL algorithms: (1) *contrastive representations*, in which methods learn "successor features" with a contrastive objective that performs inference over future outcomes, and (2) *temporal distances*, which link the (quasimetric) distance in representation space to the transit time from states to goals. We propose an approach that unifies these two frameworks, using the structure of a quasimetric representation space (triangle inequality) with the right additional constraints to learn successor representations that enable optimal goal-reaching. Unlike past work, our approach is able to exploit a **quasimetric** distance parameterization to learn **optimal** goal-reaching distances, even with **suboptimal** data and in **stochastic** environments. This gives us the best of both worlds: we retain the stability and long-horizon capabilities of Monte Carlo contrastive RL methods, while getting the free stitching capabilities of quasimetric network parameterizations. On existing offline GCRL benchmarks, our representation learning objective improves performance on stitching tasks where methods based on contrastive learning struggle, and on noisy, high-dimensional environments where methods based on quasimetric networks struggle.

强化学习目标导向表示学习离线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。