用分布式批评者提升流模型策略的离线强化学习性能
Unleashing Flow Policies with Distributional Critics
- 用流匹配建模回报分布,替代单一值估计
- 在D4RL和OGBench上显著优于现有方法
- 特别适合需要多模态动作的行为学习任务
基于流的策略近期在离线及离线到在线强化学习中表现出强大能力,可建模预收集数据集中复杂的多模态行为。然而,这些表达力强的策略常受限于其批评者——通常仅学习期望回报的单一标量估计。为解决此问题,我们提出分布式流批评者(DFC),一种新型批评者架构,可学习状态-动作回报的完整分布。与回归单一值不同,DFC采用流匹配技术,将简单基分布连续、灵活地变换为目标回报分布。由此,DFC为表达力强的流策略提供丰富且分布式的贝尔曼目标,带来更稳定、更丰富的学习信号。在D4RL和OGBench基准上的大量实验表明,该方法在需多模态动作分布的任务上表现优异,且在离线及离线转在线微调中均优于现有方法。
原文摘要 · Abstract (English)
Flow-based policies have recently emerged as a powerful tool in offline and offline-to-online reinforcement learning, capable of modeling the complex, multimodal behaviors found in pre-collected datasets. However, the full potential of these expressive actors is often bottlenecked by their critics, which typically learn a single, scalar estimate of the expected return. To address this limitation, we introduce the Distributional Flow Critic (DFC), a novel critic architecture that learns the complete state-action return distribution. Instead of regressing to a single value, DFC employs flow matching to model the distribution of return as a continuous, flexible transformation from a simple base distribution to the complex target distribution of returns. By doing so, DFC provides the expressive flow-based policy with a rich, distributional Bellman target, which offers a more stable and informative learning signal. Extensive experiments across D4RL and OGBench benchmarks demonstrate that our approach achieves strong performance, especially on tasks requiring multimodal action distributions, and excels in both offline and offline-to-online fine-tuning compared to existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。