arXiv:2606.05468cs.RO2026-06被引 7

无需奖励信号,通过对比优化让视觉语言动作模型更可靠地控制机器人

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization

论文配图:FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization
图 1 · 摘自论文原文
  • 用对比+正则化方法优化动作头,避免奖励作弊
  • 真实机器人操作生成正负轨迹对,提升训练效率
  • 适合想在不设计奖励的情况下提升机器人策略的开发者

将视觉语言动作(VLA)模型微调为可部署于真实机器人的策略仍是主要瓶颈。SFT和DAgger仅间接利用失败信号,而基于奖励的强化学习受限于现实奖励设计困难及可靠价值函数训练难题。本文提出FlowPRO,一种针对流匹配型VLA的免奖励离线强化微调框架。算法上,提出RPRO(机器人流匹配近端偏好优化),专为VLA的动作头设计的偏好优化目标,结合对比优化器与显式近端正则项,锚定隐式奖励的绝对大小,从而消除纯Flow-DPO中的奖励作弊问题。数据层面,采用遥操作干预回滚范式,从单次操作生成自然配对的正负轨迹(τ^w, τ^l);再通过平滑插值与批量混合,将稀疏修正转化为密集的状态级监督,同时保留基础策略能力。在四个长时程双臂任务中,FlowPRO达到最高成功率,超越四种代表性基线,消融实验验证了各损失组件的有效性。

原文摘要 · Abstract (English)

Post-training Vision-Language-Action (VLA) models into policies that can be reliably deployed on real robots remains a major bottleneck. SFT and DAgger exploit failure signals only indirectly, and reward-based RL is bottlenecked by the difficulty of real-world reward design and of training reliable critics. We present FlowPRO, a reward-free offline reinforced fine-tuning framework for flow-matching VLAs. Algorithmically, we propose RPRO (Robotic Flow-matching Proximalized Preference Optimization), a preference-optimization objective tailored to the flow-matching action head of VLA models. RPRO pairs a contrastive optimizer with an explicit proximal regularizer that anchors the absolute magnitude of the implicit reward, thereby eliminating the reward-hacking failure mode of plain Flow-DPO. On the data side, a teleoperated intervention-and-rollback paradigm produces naturally paired positive and negative trajectories $(τ^w, τ^l)$ on a real robot from a single operator action; a Smooth Interpolation procedure, combined with batch mixing, then converts these sparse corrections into dense per-state supervision while preserving the base policy's capabilities. On four long-horizon bimanual tasks, FlowPRO attains the highest success rate, outperforming four representative baselines, and ablations confirm the contribution of each loss component.

机器人控制免奖励学习流匹配偏好优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。