用流模型统一建模多模态动作与策略优化,提升机器人控制性能。
Decision Flow Policy Optimization
- 将流模型的动作生成视为决策过程,实现分布建模与策略优化同步。
- 在数十个离线强化学习环境中达到或超越现有最优表现。
- 适合研究多模态决策、机器人控制与生成式强化学习的学者。
近年来,生成模型在图像、视频、语言和决策等多个领域展现出卓越能力。通过将基于流的生成模型应用于强化学习,可有效建模复杂的多模态动作分布,在连续动作空间中实现优于传统高斯策略的机器人控制。以往方法通常将生成模型作为行为模型,从数据集中拟合状态条件下的动作分布,而策略优化则通过额外的值基采样加权或梯度更新独立进行。这种分离机制导致多模态分布拟合与策略改进无法同时优化,限制了模型训练并影响性能。为此,本文提出决策流(Decision Flow),一个统一框架,整合多模态动作分布建模与策略优化。具体而言,我们将基于流的模型的动作生成过程形式化为流决策过程,每个动作生成步骤对应一次流决策,从而无缝地在捕捉多模态动作分布的同时优化流策略。我们提供了决策流的严格理论证明,并在数十个离线强化学习环境上进行了广泛实验验证。结果表明,相较于现有离线强化学习基准方法,本方法实现了或达到了最先进性能。
原文摘要 · Abstract (English)
In recent years, generative models have shown remarkable capabilities across diverse fields, including images, videos, language, and decision-making. By applying powerful generative models such as flow-based models to reinforcement learning, we can effectively model complex multi-modal action distributions and achieve superior robotic control in continuous action spaces, surpassing the limitations of single-modal action distributions with traditional Gaussian-based policies. Previous methods usually adopt the generative models as behavior models to fit state-conditioned action distributions from datasets, with policy optimization conducted separately through additional policies using value-based sample weighting or gradient-based updates. However, this separation prevents the simultaneous optimization of multi-modal distribution fitting and policy improvement, ultimately hindering the training of models and degrading the performance. To address this issue, we propose Decision Flow, a unified framework that integrates multi-modal action distribution modeling and policy optimization. Specifically, our method formulates the action generation procedure of flow-based models as a flow decision-making process, where each action generation step corresponds to one flow decision. Consequently, our method seamlessly optimizes the flow policy while capturing multi-modal action distributions. We provide rigorous proofs of Decision Flow and validate the effectiveness through extensive experiments across dozens of offline RL environments. Compared with established offline RL baselines, the results demonstrate that our method achieves or matches the SOTA performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。