arXiv:2506.00795cs.LG2025-06

让监督学习具备轨迹拼接能力,缩小与强化学习的性能差距

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization

  • 用Q函数条件化策略和最大化,增强监督学习的轨迹拼接能力
  • 在离线数据集上达到优于现有监督学习方法的性能
  • 适合追求稳定高效且需轨迹推理的离线目标导向强化学习场景

近期,监督学习(SL)因其简单、稳定和高效,成为离线强化学习的有效方法。然而,已有研究指出,SL方法缺乏通常由时序差分(TD)方法具有的轨迹拼接能力。一个自然的问题是:如何使SL方法具备拼接能力,并缩小其与TD学习的性能差距?为此,本文提出基于Q函数条件化的监督学习方法,用于离线目标条件化强化学习,通过引入Q条件化策略和Q条件化最大化,赋予SL拼接能力。具体地,我们提出目标条件化强化监督学习(GCReinSL),包括:(1) 利用归一化流从离线数据集中估计Q函数;(2) 通过将Q函数最大化与期望值回归结合,在数据支持范围内寻找最大Q值。推理时,策略基于该最大Q值选择最优动作。在多个离线强化学习数据集上的拼接评估实验表明,本方法在具备拼接能力的同时,显著优于先前具有拼接能力和目标数据增强技术的监督学习方法。

原文摘要 · Abstract (English)

Recently, supervised learning (SL) methodology has emerged as an effective approach for offline reinforcement learning (RL) due to their simplicity, stability, and efficiency. However, recent studies show that SL methods lack the trajectory stitching capability, typically associated with temporal difference (TD)-based approaches. A question naturally surfaces: \textit{How can we endow SL methods with stitching capability and close its performance gap with TD learning?} To answer this question, we introduce $Q$-conditioned maximization supervised learning for offline goal-conditioned RL, which enhances SL with the stitching capability through $Q$-conditioned policy and $Q$-conditioned maximization. Concretely, we propose \textbf{G}oal-\textbf{C}onditioned \textbf{\textit{Rein}}forced \textbf{S}upervised \textbf{L}earning (\textbf{GC\textit{Rein}SL}), which consists of (1) estimating the $Q$-function by Normalizing Flows from the offline dataset and (2) finding the maximum $Q$-value within the data support by integrating $Q$-function maximization with Expectile Regression. In inference time, our policy chooses optimal actions based on such a maximum $Q$-value. Experimental results from stitching evaluations on offline RL datasets demonstrate that our method outperforms prior SL approaches with stitching capabilities and goal data augmentation techniques.

强化学习监督学习离线学习轨迹拼接

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。