arXiv:2605.27079cs.LGcs.AI2026-05被引 2

提出TRQAM算法,稳定微调预训练流策略的离线强化学习。

Trust Region Q Adjoint Matching

论文配图:Trust Region Q Adjoint Matching
图 1 · 摘自论文原文
  • 通过投影对偶下降自适应控制路径空间KL散度
  • 在50个OGBench任务上离线强化学习成功率达68%(基线46%)
  • 适合需要稳定微调预训练策略的研究者

预训练流策略的离线强化学习因多步采样导致优化不稳定而难以实现。近期提出的Q-伴随匹配(QAM)将其重构为无记忆随机最优控制(SOC)问题,并引入学习的评判器解决此问题。然而,QAM继承了评判器引导改进的根本缺陷:当评判器病态时,小误差会被放大,常引发模型坍塌。本文提出信任域Q-伴随匹配(TRQAM),一种稳定的离线微调算法,通过投影对偶下降自适应控制预训练流策略的路径空间KL散度。具体地,我们优化SOC动力学中的信任域参数λ,理论证明路径空间KL可表示为λ的闭式函数。由此,方法能精确控制与预训练策略的偏差,实现稳定离线强化学习。在50个OGBench任务上的实验表明,TRQAM在离线强化学习和离线到在线强化学习中均持续优于现有方法,尤其在离线强化学习中取得68%的整体成功率,显著超越最强基线(46%)。

原文摘要 · Abstract (English)

Off-policy reinforcement learning of pretrained flow policies remains challenging due to the instability of optimization arising from the multi-step sampling process. Recently, Q-learning with Adjoint Matching (QAM) addressed this issue by reformulating into a memoryless stochastic optimal control (SOC) problem with a learned critic. However, QAM inherits a fundamental fragility of critic-guided improvement: small critic errors are amplified when critics are ill-conditioned, often leading to model collapse. This paper introduces Trust Region Q-Adjoint Matching (TRQAM), a stable off-policy fine-tuning algorithm that adaptively controls the path-space KL with pretrained flow policies through projected dual descent. Specifically, we optimize the trust-region parameter $λ$ in SOC dynamics, and theoretically show that the path-space KL can be represented by a closed-form function of $λ$. As a result, our method can precisely control the exact deviation from pretrained flow policies, achieving stable off-policy RL. Through experiments on 50 OGBench tasks, TRQAM consistently outperforms prior arts in both offline RL and offline-to-online RL. In particular, TRQAM achieves an overall success rate of 68% in offline RL, substantially improves the strongest baseline at 46%.

强化学习离线学习流模型稳定训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。