arXiv:2604.10962cs.RO2026-04被引 1

用得分函数调节轨迹漂移,实现更高效的机器人控制优化。

ScoRe-Flow: Complete Distributional Control via Score-Based Reinforcement Learning for Flow Matching

  • 通过得分函数调控漂移,引导探索向高概率区域集中。
  • 在D4RL任务上收敛速度提升2.4倍,操纵任务成功率最高提高5.4%。
  • 无需额外网络,适合需要稳定高效强化学习微调的机器人系统。

流匹配(Flow Matching, FM)策略已成为机器人控制的高效基础,能快速生成丰富动作,支撑当前大规模具身AI系统。然而,基于模仿学习训练的FM策略继承了示范数据的局限性;要超越次优行为,需进行强化学习(RL)微调。现有方法将确定性流转化为带可学习噪声注入的随机微分方程(SDE),虽支持探索和可计算似然,但仅靠噪声控制会削弱训练效率,尤其当示范数据已提供强先验时。我们观察到,通过得分函数(即对数密度梯度)调节漂移项,可引导探索朝高概率区域推进,提升稳定性。该得分函数可从速度场直接获得闭式表达,无需额外网络。基于此,我们提出ScoRe-Flow,一种基于得分的强化学习微调方法,结合漂移调制与可学习方差预测,实现对随机转移均值与方差的解耦控制。实验表明,ScoRe-Flow在D4RL运动任务上收敛速度比现有最优方法快2.4倍,在Robomimic与Franka Kitchen操纵任务上成功率最高提升5.4%。

原文摘要 · Abstract (English)

Flow Matching (FM) policies have emerged as an efficient backbone for robotic control, offering fast and expressive action generation that underpins recent large-scale embodied AI systems. However, FM policies trained via imitation learning inherit the limitations of demonstration data; surpassing suboptimal behaviors requires reinforcement learning (RL) fine-tuning. Recent methods convert deterministic flows into stochastic differential equations (SDEs) with learnable noise injection, enabling exploration and tractable likelihoods, but such noise-only control can compromise training efficiency when demonstrations already provide strong priors. We observe that modulating the drift via the score function, i.e., the gradient of log-density, steers exploration toward high-probability regions, improving stability. The score admits a closed-form expression from the velocity field, requiring no auxiliary networks. Based on this, we propose ScoRe-Flow, a score-based RL fine-tuning method that combines drift modulation with learned variance prediction to achieve decoupled control over the mean and variance of stochastic transitions. Experiments demonstrate that ScoRe-Flow achieves 2.4x faster convergence than flow-based SOTA on D4RL locomotion tasks and up to 5.4% higher success rates on Robomimic and Franka Kitchen manipulation tasks.

机器人控制流匹配强化学习得分函数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。