arXiv:2606.29820cs.LGcs.AI2026-06

用双流框架提升复杂控制任务中的探索与价值估计精度

Dual-Flow Reinforcement Learning with State-Aware Exploration

论文配图:Dual-Flow Reinforcement Learning with State-Aware Exploration
图 1 · 摘自论文原文
  • 通过条件流匹配联合建模回报与动作的多模态分布
  • 在多个基准上超越现有扩散与流模型,性能领先
  • 引入状态感知探索调节器,避免策略坍缩并覆盖高价值区域

在复杂的连续控制强化学习任务中,最优动作常呈现多模态特性,且对应回报分布也高度不确定,导致价值估计与多模态探索困难。现有使用单峰高斯的价值估计方法表达能力有限,易产生偏差;近期生成式策略虽可表示多模态动作,但常坍缩至少数模式,且未能充分探索高价值动作空间。为此,我们提出 Dual-Flow RL,一个统一的演员-评论家框架,利用条件流匹配(CFM)同时建模连续回报分布与多模态策略分布,实现可靠的值估计与持续的多模态探索。为进一步增强探索,我们设计了熵-协方差探索调节器(ECER),通过策略熵与动作不确定性协方差实现状态感知的探索调控。在 DeepMind Control Suite 与 Humanoid-Bench 上的实验表明,Dual-Flow RL 在多数任务上达到当前最优性能,显著优于先前基于扩散与流的方法。

原文摘要 · Abstract (English)

In complex continuous-control reinforcement learning tasks, multimodal optimal actions often coincide with uncertain, multimodal return distributions, making reliable value estimation and multimodal exploration challenging. Existing value estimation methods using unimodal Gaussians restrict expressiveness and yield biased estimates. Recent generative policies can represent multimodal actions but often collapse to a few modes and under-explore high-value areas of the action space. Motivated by these challenges, we propose Dual-Flow RL, a unified actor-critic framework that jointly models a continuous return distribution and a multimodal policy distribution using conditional flow matching (CFM). This design supports reliable value estimation and sustained multimodal exploration. To further enhance exploration, we introduce an Entropy-Covariance Exploration Regulator (ECER) that enables state-aware exploration regulation leveraging policy entropy and action-uncertainty covariance. Experiments on DeepMind Control Suite and Humanoid-Bench show that Dual-Flow RL achieves state-of-the-art performance on most tasks, significantly outperforming prior diffusion-based and flow-based methods.

强化学习多模态探索流模型价值估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。