arXiv:2602.18015cs.LGcs.AI2026-02中稿 · ICLR被引 6

用流模型提升离线强化学习的策略与价值函数,防止数据外过估计。

Flow Actor-Critic for Offline Reinforcement Learning

  • 结合流模型设计演员与保守评论家,联合优化策略表达能力。
  • 在D4RL和OGBench上达到新最佳性能,有效抑制数据外Q值爆炸。
  • 适合研究复杂多模态决策问题的离线强化学习方向。

离线强化学习中的数据分布常呈现复杂且多模态特性,需要更强大的策略来捕捉这些分布,超越传统高斯策略。本文提出Flow Actor-Critic,一种基于最新流策略的新型离线强化学习演员-评论家方法。该方法不仅在演员部分使用流模型,还利用流模型构建保守评论家,以防止数据外区域的Q值爆炸。为此,我们提出一种新的评论家正则化项,基于流策略设计中产生的行为代理模型。通过这种联合机制,该方法在D4RL及近期OGBench基准测试中均取得新最佳性能。

原文摘要 · Abstract (English)

The dataset distributions in offline reinforcement learning (RL) often exhibit complex and multi-modal distributions, necessitating expressive policies to capture such distributions beyond widely-used Gaussian policies. To handle such complex and multi-modal datasets, in this paper, we propose Flow Actor-Critic, a new actor-critic method for offline RL, based on recent flow policies. The proposed method not only uses the flow model for actor as in previous flow policies but also exploits the expressive flow model for conservative critic acquisition to prevent Q-value explosion in out-of-data regions. To this end, we propose a new form of critic regularizer based on the flow behavior proxy model obtained as a byproduct of flow-based actor design. Leveraging the flow model in this joint way, we achieve new state-of-the-art performance for test datasets of offline RL including the D4RL and recent OGBench benchmarks.

离线RL流模型策略优化价值函数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。