让视觉语言动作模型通过真实任务经验自我改进,提升执行效率与成功率。
$π^{*}_{0.6}$: a VLA That Learns From Experience

- 基于优势条件策略的强化学习方法,融合多种真实数据实现自优化。
- 在复杂家务任务中,任务完成率提升一倍以上,失败率降低近半。
- 适合希望提升机器人自主执行能力的研究者和工程团队。
我们研究了视觉-语言-动作(VLA)模型如何通过真实世界部署实现强化学习(RL)改进。提出一种通用方法——带经验与修正的优势条件策略强化学习(RECAP),通过优势条件化实现VLA的强化学习训练。该方法整合了示范数据、在线采集数据及自主执行中的专家远程干预数据。首先使用离线强化学习预训练一个通用型VLA,称为π*_{0.6},随后通过机器人端数据收集进行任务特化。实验表明,采用完整RECAP方法训练的π*_{0.6}模型可在真实家庭环境中可靠折叠衣物、组装盒子,并使用专业咖啡机制作意式浓缩咖啡。在部分最困难的任务中,RECAP使任务吞吐量提升超过一倍,任务失败率降低约一半。
原文摘要 · Abstract (English)
We study how vision-language-action (VLA) models can improve through real-world deployments via reinforcement learning (RL). We present a general-purpose method, RL with Experience and Corrections via Advantage-conditioned Policies (RECAP), that provides for RL training of VLAs via advantage conditioning. Our method incorporates heterogeneous data into the self-improvement process, including demonstrations, data from on-policy collection, and expert teleoperated interventions provided during autonomous execution. RECAP starts by pre-training a generalist VLA with offline RL, which we call $π^{*}_{0.6}$, that can then be specialized to attain high performance on downstream tasks through on-robot data collection. We show that the $π^{*}_{0.6}$ model trained with the full RECAP method can fold laundry in real homes, reliably assemble boxes, and make espresso drinks using a professional espresso machine. On some of the hardest tasks, RECAP more than doubles task throughput and roughly halves the task failure rate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。