arXiv:2602.02481cs.ROcs.AI2026-02被引 9

用流匹配方法训练机器人控制策略,提升复杂任务表现与迁移能力。

Flow Policy Gradients for Robot Control

  • 采用流匹配框架绕过似然计算,支持更复杂的动作分布建模。
  • 在足式行走、人形追踪和操作任务中实现成功训练,且模拟到现实迁移稳定。
  • 新目标函数增强探索能力,微调时比基线更鲁棒,适合高精度控制场景。

基于似然的策略梯度方法是当前从奖励信号训练机器人控制策略的主流方式,但依赖可微分的动作似然,限制了策略输出为高斯等简单分布。本文展示如何将近期无需似然计算的流匹配策略梯度框架有效应用于挑战性机器人控制任务中,以训练更具表达力的策略。我们提出一种改进目标函数,在足式行走、人形运动追踪及操作任务中均取得成功,并在两台人形机器人上实现了稳健的模拟到现实迁移。通过消融实验与训练动态分析,结果表明:从零开始训练时,策略能利用流表示实现更优探索;微调阶段也显著优于基线方法。

原文摘要 · Abstract (English)

Likelihood-based policy gradient methods are the dominant approach for training robot control policies from rewards. These methods rely on differentiable action likelihoods, which constrain policy outputs to simple distributions like Gaussians. In this work, we show how flow matching policy gradients -- a recent framework that bypasses likelihood computation -- can be made effective for training and fine-tuning more expressive policies in challenging robot control settings. We introduce an improved objective that enables success in legged locomotion, humanoid motion tracking, and manipulation tasks, as well as robust sim-to-real transfer on two humanoid robots. We then present ablations and analysis on training dynamics. Results show how policies can exploit the flow representation for exploration when training from scratch, as well as improved fine-tuning robustness over baselines.

机器人控制流匹配策略梯度强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。