arXiv:2509.06782cs.LG2025-09NeurIPS被引 13

用物理规律约束价值函数,提升离线目标导向强化学习的泛化能力。

Physics-informed Value Learner for Offline Goal-Conditioned Reinforcement Learning

  • 基于欧几里得偏微分方程设计物理感知正则化项,引导价值函数符合最优控制结构。
  • 在长时序任务和大规模导航中性能显著提升,尤其在拼接区域表现突出。
  • 可无缝集成到现有离线目标导向强化学习算法,适合复杂环境下的安全智能体训练。

离线目标导向强化学习(Offline GCRL)在自主导航与运动控制等场景中具有巨大潜力,但受限于数据集对状态-动作空间覆盖不足,以及长时序任务的泛化能力。为此,本文提出一种基于欧几里得偏微分方程(Eikonal PDE)的物理感知(Pi)正则化损失,为价值学习引入几何归纳偏置。该正则化项源自连续时间最优控制理论,不依赖通用梯度惩罚,而是促使价值函数与代价到目标(cost-to-go)结构对齐。该方法与基于时序差分的价值学习兼容,可集成至现有离线GCRL算法中。结合分层隐式Q学习(HIQL),提出的欧几里得正则化HIQL(Eik-HIQL)在性能与泛化性上均有显著提升,尤其在拼接区域与大规模导航任务中优势明显。

原文摘要 · Abstract (English)

Offline Goal-Conditioned Reinforcement Learning (GCRL) holds great promise for domains such as autonomous navigation and locomotion, where collecting interactive data is costly and unsafe. However, it remains challenging in practice due to the need to learn from datasets with limited coverage of the state-action space and to generalize across long-horizon tasks. To improve on these challenges, we propose a \emph{Physics-informed (Pi)} regularized loss for value learning, derived from the Eikonal Partial Differential Equation (PDE) and which induces a geometric inductive bias in the learned value function. Unlike generic gradient penalties that are primarily used to stabilize training, our formulation is grounded in continuous-time optimal control and encourages value functions to align with cost-to-go structures. The proposed regularizer is broadly compatible with temporal-difference-based value learning and can be integrated into existing Offline GCRL algorithms. When combined with Hierarchical Implicit Q-Learning (HIQL), the resulting method, Eikonal-regularized HIQL (Eik-HIQL), yields significant improvements in both performance and generalization, with pronounced gains in stitching regimes and large-scale navigation tasks.

强化学习离线学习物理先验导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。