让机器人通过交互引导的强化学习,更可靠地完成复杂长程操作任务。
IG-RFT: An Interaction-Guided RL Framework for VLA Models in Long-Horizon Robotic Manipulation
- 根据机器人交互状态动态调节探索强度,提升训练效率。
- 混合轨迹与子任务奖励,解决稀疏奖励问题,成功率提升至85%。
- 三阶段强化学习流程,适合真实世界长程机器人任务优化。
视觉-语言-动作(VLA)模型在通用机器人策略中展现出巨大潜力,但在新颖真实场景中处理长时程复杂任务时,受限于分布偏移和高质量示范数据稀缺,泛化能力不足。尽管强化学习(RL)为策略改进提供了可能,但其在真实世界VLA微调中面临探索效率低、训练不稳定和样本成本高等挑战。为此,我们提出IG-RFT,一种专为基于流的VLA模型设计的交互引导强化微调系统。首先,引入交互引导的优势加权回归(IG-AWR)算法,根据机器人交互状态动态调节探索强度。其次,设计融合轨迹级与子任务级奖励的混合密集奖励函数,缓解稀疏或任务特定奖励的问题。最后,构建包含SFT、离线RL和人机协同RL的三阶段强化学习系统以微调VLA模型。在四个具有挑战性的长时程任务上的真实世界实验表明,IG-RFT平均成功率达85.0%,显著优于SFT(18.8%)和标准离线RL基线(40.0%)。消融实验验证了IG-AWR和混合奖励设计的关键作用。本工作建立并验证了面向真实世界机器人操作的新型强化微调系统。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have demonstrated significant potential for generalist robotic policies; however, they struggle to generalize to long-horizon complex tasks in novel real-world domains due to distribution shifts and the scarcity of high-quality demonstrations. Although reinforcement learning (RL) offers a promising avenue for policy improvement, applying it to real-world VLA fine-tuning faces challenges regarding exploration efficiency, training stability, and sample cost. To address these issues, we propose IG-RFT, a novel Interaction-Guided Reinforced Fine-Tuning system designed for flow-based VLA models. Firstly, to facilitate effective policy optimization, we introduce Interaction-Guided Advantage Weighted Regression (IG-AWR), an RL algorithm that dynamically modulates exploration intensity based on the robot's interaction status. Furthermore, to address the limitations of sparse or task-specific rewards, we design a novel hybrid dense reward function that integrates the trajectory-level reward and the subtask-level reward. Finally, we construct a three-stage RL system comprising SFT, Offline RL, and Human-in-the-Loop RL for fine-tuning VLA models. Extensive real-world experiments on four challenging long-horizon tasks demonstrate that IG-RFT achieves an average success rate of 85.0%, significantly outperforming SFT (18.8%) and standard Offline RL baselines (40.0%). Ablation studies confirm the critical contributions of IG-AWR and hybrid reward shaping. In summary, our work establishes and validates a novel reinforced fine-tuning system for VLA models in real-world robotic manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。