提升离线强化学习中复杂行为的探索能力,避免低密度区域高奖励动作被忽略。
Entropy-Regularized Adjoint Matching for Offline Reinforcement Learning

- 用熵正则化改进连续伴随方法,缓解行为分布偏好问题。
- 引入混合行为先验,扩大探索范围至分布外高回报区域。
- 适合需要安全高效探索的复杂连续控制任务研究者。
将表达性强的生成策略(如流匹配模型)融入离线强化学习,使智能体能捕捉复杂、多模态行为。尽管基于伴随匹配的Q-learning(QAM)通过连续伴随方法稳定了策略优化,但其仍受限于固定的观测行为分布,导致存在‘流行度偏差’,抑制低密度区域的高奖励动作,并引发‘支持绑定’,限制分布外探索。现有方法如添加残差高斯策略,常重新引入单峰分布的表达瓶颈。本文提出最大熵伴随匹配(ME-AM),在连续流框架下统一解决上述问题:(1) 采用镜面下降熵最大化目标,缓解流行度偏差,从离线数据集中提取最优策略;(2) 引入混合行为先验,扩展几何支持以涵盖分布外高回报区域。通过探索该扩展几何空间,ME-AM识别出鲁棒动作,同时保持生成向量场的绝对连续性。实验表明,在多种稀疏奖励连续控制环境中,ME-AM性能优于或媲美现有最先进方法。
原文摘要 · Abstract (English)
Integrating expressive generative policies, such as flow-matching models, into offline reinforcement learning (RL) allows agents to capture complex, multi-modal behaviors. While Q-learning with Adjoint Matching (QAM) stabilizes policy optimization via the continuous adjoint method, it remains inherently bound to the fixed behavior distribution. This dependence induces a \textit{popularity bias} that can suppress high-reward actions in low-density regions, and creates a \textit{support binding} that restricts off-manifold exploration. Existing workarounds, such as appending \textit{residual} Gaussian policies, often re-introduce the expressivity bottlenecks associated with unimodal distributions. In this work, we propose \textit{Maximum Entropy Adjoint Matching} (ME-AM), a unified framework that addresses these limitations within the continuous flow formulation. ME-AM incorporates two mechanisms: (1) a Mirror Descent entropy maximization objective that mitigates the popularity bias to facilitate the extraction of optimal policies from offline datasets, and (2) a \textit{Mixture Behavior Prior} that broadens the geometric support to encompass out-of-distribution high-reward regions. By exploring this extended geometry, ME-AM identifies robust actions while preserving the absolute continuity of the generative vector field. Empirically, ME-AM demonstrates competitive or superior performance compared to prior state-of-the-art (SOTA) methods across a diverse suite of sparse-reward continuous control environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。