arXiv:2507.15073cs.LG2025-07被引 15

用强化学习提升机器人动作生成模型性能,显著优于人类示范。

Reinforcement Learning for Flow-Matching Policies

  • 通过强化学习优化流匹配策略,实现更优控制
  • GRPO方法比基础模仿学习降低50%~85%成本
  • 适合需要高效运动规划的机器人任务

流匹配策略已成为通用机器人领域的有力范式,其通过条件化传感器观测与文本指令来模仿动作片段。通常,训练数据由次优策略(如人工操作)生成。本文探索通过强化学习训练流匹配策略,以超越原始示范策略的表现。特别关注最短时间控制这一关键应用场景,并提出一种可变时域流匹配规划方案。随后引入两类方法:简单的奖励加权流匹配(RWFM)和基于学习奖励代理的组相对策略优化(GRPO)。模型在模拟单轮车动力学任务上进行训练,结果表明两种方法均显著优于次优示范者表现,其中GRPO方法相较朴素模仿学习流匹配(ILFM)通常降低成本50%至85%。

原文摘要 · Abstract (English)

Flow-matching policies have emerged as a powerful paradigm for generalist robotics. These models are trained to imitate an action chunk, conditioned on sensor observations and textual instructions. Often, training demonstrations are generated by a suboptimal policy, such as a human operator. This work explores training flow-matching policies via reinforcement learning to surpass the original demonstration policy performance. We particularly note minimum-time control as a key application and present a simple scheme for variable-horizon flow-matching planning. We then introduce two families of approaches: a simple Reward-Weighted Flow Matching (RWFM) scheme and a Group Relative Policy Optimization (GRPO) approach with a learned reward surrogate. Our policies are trained on an illustrative suite of simulated unicycle dynamics tasks, and we show that both approaches dramatically improve upon the suboptimal demonstrator performance, with the GRPO approach in particular generally incurring between $50\%$ and $85\%$ less cost than a naive Imitation Learning Flow Matching (ILFM) approach.

强化学习机器人控制流匹配最优控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。