用几何距离规划长程目标导航,提升离线强化学习的稳定性与效率。
Offline Goal-Conditioned Reinforcement Learning with Projective Quasimetric Planning
- 构建非对称距离度量,引导稀疏关键点均匀分布于隐空间。
- 通过方向性代价函数实现子目标渐进逼近,减少误差累积。
- 适合复杂长程导航任务,尤其在数据有限场景下表现稳健。
离线目标条件强化学习旨在利用先前收集的轨迹训练智能体完成指定目标。然而,将其拓展至长时序任务仍面临挑战,主要源于累积的价值估计误差。受几何建模的启发,我们提出投影拟度量规划(ProQ),一种组合式框架:首先学习一个非对称距离度量,将其用于两个目的——一是作为排斥能量,促使一组稀疏关键点在学习到的隐空间中均匀分布;二是作为结构化方向性成本,指导向邻近子目标前进。特别地,ProQ将该几何结构与拉格朗日异常检测器结合,确保学习的关键点始终位于可到达区域内。通过统一度量学习、关键点覆盖与目标条件控制,该方法生成有意义的子目标,并在多个导航基准上实现了鲁棒的长程目标达成能力。
原文摘要 · Abstract (English)
Offline Goal-Conditioned Reinforcement Learning seeks to train agents to reach specified goals from previously collected trajectories. Scaling that promises to long-horizon tasks remains challenging, notably due to compounding value-estimation errors. Principled geometric offers a potential solution to address these issues. Following this insight, we introduce Projective Quasimetric Planning (ProQ), a compositional framework that learns an asymmetric distance and then repurposes it, firstly as a repulsive energy forcing a sparse set of keypoints to uniformly spread over the learned latent space, and secondly as a structured directional cost guiding towards proximal sub-goals. In particular, ProQ couples this geometry with a Lagrangian out-of-distribution detector to ensure the learned keypoints stay within reachable areas. By unifying metric learning, keypoint coverage, and goal-conditioned control, our approach produces meaningful sub-goals and robustly drives long-horizon goal-reaching on diverse a navigation benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。