改进视觉语言动作模型的离线强化学习微调,提升复杂任务表现。
Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models
- 引入自适应缩放因子,平衡强化学习信号与梯度方差。
- 在仿真和真实场景中均实现强泛化与少样本学习能力。
- 适合需要稳定微调的机器人操作研究者使用。
基于流匹配的视觉-语言-动作(VLA)模型在通用机器人操作任务中表现优异,但在复杂下游任务上的动作精度仍不理想。主要原因在于其仅依赖模仿学习的后训练范式,难以深入理解数据质量分布特性,而强化学习(RL)正擅长此点。本文理论提出一种适用于VLA流模型的离线强化学习后训练目标,并设计高效可行的微调算法——自适应增强流匹配(ARFM)。通过在VLA流模型损失中引入自适应调整的缩放因子,构建了具有原则性偏差-方差权衡的目标函数,以最优控制强化学习信号对流损失的影响。ARFM 自适应地平衡强化学习优势保留与流损失梯度方差控制,实现更稳定高效的微调过程。大量仿真与真实世界实验表明,ARFM 在泛化性、鲁棒性、少样本学习及持续学习方面均表现优异。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models based on flow matching have shown excellent performance in general-purpose robotic manipulation tasks. However, the action accuracy of these models on complex downstream tasks is unsatisfactory. One important reason is that these models rely solely on the post-training paradigm of imitation learning, which makes it difficult to have a deeper understanding of the distribution properties of data quality, which is exactly what Reinforcement Learning (RL) excels at. In this paper, we theoretically propose an offline RL post-training objective for VLA flow models and induce an efficient and feasible offline RL fine-tuning algorithm -- Adaptive Reinforced Flow Matching (ARFM). By introducing an adaptively adjusted scaling factor in the VLA flow model loss, we construct a principled bias-variance trade-off objective function to optimally control the impact of RL signal on flow loss. ARFM adaptively balances RL advantage preservation and flow loss gradient variance control, resulting in a more stable and efficient fine-tuning process. Extensive simulation and real-world experimental results show that ARFM exhibits excellent generalization, robustness, few-shot learning, and continuous learning performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。