arXiv:2512.12046cs.LGcs.RO2025-12被引 6

用偏微分方程实现无需轨迹的智能体目标导航,提升泛化能力。

Goal Reaching with Eikonal-Constrained Hierarchical Quasimetric Reinforcement Learning

  • 基于欧几里得方程构建连续时间值函数,无需依赖完整轨迹。
  • 在离线导航任务中达到当前最优性能,且在操作任务中稳定优于旧方法。
  • 适合复杂动态环境下的目标导向强化学习研究者使用。

目标条件强化学习(GCRL)通过将任务定义为到达目标而非最大化人工奖励信号,缓解了奖励设计难题。在此设定下,最优目标条件价值函数自然形成拟度量,推动了拟度量强化学习(QRL)的发展,其通过约束价值学习为拟度量映射,并利用离散轨迹基约束保证局部一致性。本文提出基于欧几里得偏微分方程(Eikonal PDE)的连续时间重构方法——Eik-QRL,使该方法无需轨迹信息,仅需采样状态与目标即可训练,同时提升分布外泛化能力。我们为Eik-QRL提供了理论保障,并识别出复杂动态下存在的局限性。为此,引入层次化结构的Eik-HiQRL,将Eik-QRL融入分层分解框架。实验表明,Eik-HiQRL在离线目标导航任务中表现领先,在操作任务中持续优于传统QRL,媲美时序差分方法。

原文摘要 · Abstract (English)

Goal-Conditioned Reinforcement Learning (GCRL) mitigates the difficulty of reward design by framing tasks as goal reaching rather than maximizing hand-crafted reward signals. In this setting, the optimal goal-conditioned value function naturally forms a quasimetric, motivating Quasimetric RL (QRL), which constrains value learning to quasimetric mappings and enforces local consistency through discrete, trajectory-based constraints. We propose Eikonal-Constrained Quasimetric RL (Eik-QRL), a continuous-time reformulation of QRL based on the Eikonal Partial Differential Equation (PDE). This PDE-based structure makes Eik-QRL trajectory-free, requiring only sampled states and goals, while improving out-of-distribution generalization. We provide theoretical guarantees for Eik-QRL and identify limitations that arise under complex dynamics. To address these challenges, we introduce Eik-Hierarchical QRL (Eik-HiQRL), which integrates Eik-QRL into a hierarchical decomposition. Empirically, Eik-HiQRL achieves state-of-the-art performance in offline goal-conditioned navigation and yields consistent gains over QRL in manipulation tasks, matching temporal-difference methods.

强化学习目标导向偏微分方程泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。