用新算法让视觉语言动作模型在线强化学习更稳定高效
Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models
- 提出FPO算法,解决流匹配模型强化学习中采样困难问题
- 在LIBERO和ALOHA上实现比基线更强的稀疏奖励学习效果
- 适合研究视觉语言动作模型在线优化与强化学习的学者
视觉-语言-动作(VLA)模型如OpenVLA、Octo和$π_0$通过大规模示范展现出强大泛化能力,但其性能仍受限于监督数据的质量与覆盖范围。强化学习(RL)为通过在线交互提升和微调VLAs提供了可能。然而,传统策略梯度方法在基于流匹配的模型中因重要性采样不可行而计算成本过高,需显式计算策略比值。为此,本文提出流策略优化(FPO)算法,通过利用条件流匹配目标的样本级变化重构重要性采样。此外,FPO通过结构感知信用分配提升梯度效率,采用截断代理目标稳定优化,多步潜在空间探索促进策略更新多样性,并引入Q-集合机制提供稳健的价值估计。我们在LIBERO基准和ALOHA仿真任务上评估FPO,对比监督学习、偏好对齐、基于扩散、自回归在线强化学习及$π_0$-FAST基线,均观察到相对于模仿先验和强竞争基线的一致提升,且在稀疏奖励下学习稳定。消融实验与潜在空间动态分析进一步验证了各组件的有效性,证明所提计算模块在在线强化学习中能实现条件流匹配目标的稳定收敛。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models such as OpenVLA, Octo, and $π_0$ have shown strong generalization by leveraging large-scale demonstrations, yet their performance is still fundamentally constrained by the quality and coverage of supervised data. Reinforcement learning (RL) provides a promising path for improving and fine-tuning VLAs through online interaction. However, conventional policy gradient methods are computationally infeasible in the context of flow-matching based models due to the intractability of the importance sampling process, which requires explicit computation of policy ratios. To overcome this limitation, we propose Flow Policy Optimization (FPO) algorithm, which reformulates importance sampling by leveraging per-sample changes in the conditional flow-matching objective. Furthermore, FPO achieves stable and scalable online reinforcement fine-tuning of the $π_0$ model by integrating structure-aware credit assignment to enhance gradient efficiency, clipped surrogate objectives to stabilize optimization, multi-step latent exploration to encourage diverse policy updates, and a Q-ensemble mechanism to provide robust value estimation. We evaluate FPO on the LIBERO benchmark and the ALOHA simulation task against supervised, preference-aligned, diffusion-based, autoregressive online RL, and $π_0$-FAST baselines, observing consistent improvements over the imitation prior and strong alternatives with stable learning under sparse rewards. In addition, ablation studies and analyses of the latent space dynamics further highlight the contributions of individual components within FPO, validating the effectiveness of the proposed computational modules and the stable convergence of the conditional flow-matching objective during online RL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。