arXiv:2604.08960cs.LG2026-04

用平均速度场提升离线目标导向强化学习的长时序控制能力

Efficient Hierarchical Implicit Flow Q-learning for Offline Goal-conditioned Reinforcement Learning

  • 引入平均速度场建模高层与低层策略,实现一步采样高效动作生成
  • 在OGBench上状态与像素任务均超越现有方法,最高提升18.3%
  • 通过LeJEPA损失增强目标表示区分度,改善泛化性能

离线目标导向强化学习(GCRL)旨在从无奖励的离线数据中学习目标条件策略。尽管已有层次化架构如HIQL取得进展,长时序控制仍受限于高斯策略表达能力不足及高层策略无法生成有效子目标。为此,本文提出目标条件均值流策略,通过学习平均速度场捕捉高低层策略的目标分布,实现一步采样高效动作生成。此外,针对目标表示不足问题,引入LeJEPA损失,在训练中排斥目标嵌入表示,增强其区分性,提升泛化能力。实验表明,该方法在OGBench基准上的状态基和像素基任务中均表现优异,显著优于现有方法。

原文摘要 · Abstract (English)

Offline goal-conditioned reinforcement learning (GCRL) is a practical reinforcement learning paradigm that aims to learn goal-conditioned policies from reward-free offline data. Despite recent advances in hierarchical architectures such as HIQL, long-horizon control in offline GCRL remains challenging due to the limited expressiveness of Gaussian policies and the inability of high-level policies to generate effective subgoals. To address these limitations, we propose the goal-conditioned mean flow policy, which introduces an average velocity field into hierarchical policy modeling for offline GCRL. Specifically, the mean flow policy captures complex target distributions for both high-level and low-level policies through a learned average velocity field, enabling efficient action generation via one-step sampling. Furthermore, considering the insufficiency of goal representation, we introduce a LeJEPA loss that repels goal representation embeddings during training, thereby encouraging more discriminative representations and improving generalization. Experimental results show that our method achieves strong performance across both state-based and pixel-based tasks in the OGBench benchmark.

强化学习离线学习目标导向层次策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。