arXiv:2605.25477cs.ROcs.AI2026-05被引 1

让视觉语言动作模型高效学新任务,30次全成功仅用19分钟。

EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models

论文配图:EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models
图 1 · 摘自论文原文
  • 基于预训练模型做强化学习微调,保留先验知识同时提升样本效率。
  • 在6项高难度操作任务中平均19.1分钟达成30次全成功,表现超越此前方法。
  • 适合希望快速部署机器人新任务的研究者与工程团队使用。

高效可靠地学习新任务是机器人领域的核心挑战。视觉-语言-动作(VLA)模型在多种操作任务中展现出强大泛化能力,但预训练策略仍难以满足实际部署所需的可靠性。强化学习(RL)微调有望弥合这一差距,但现有方法或从零训练未充分利用预训练先验,或微调时样本效率与成功率不足。本文提出EXPO-FT系统,实现稳定且样本高效的VLA策略强化学习微调,成功完成多项复杂操作任务,包括穿灯串、击球入袋、插花入瓶等,均需高精度、动态动作及对初始状态的鲁棒性。系统在所有测试任务中均达到30/30的成功率,平均仅需19.1分钟在线机器人数据,显著优于先前的从零训练和VLA微调方法。我们开源了代码库,以促进该技术在机器人领域更广泛应用。

原文摘要 · Abstract (English)

The ability to efficiently and reliably learn new tasks has been a foundational challenge in robotics. Vision-Language-Action (VLA) models have demonstrated strong generalization across diverse manipulation tasks, yet pretrained policies consistently fall short of the reliability required for real-world deployment. Reinforcement learning (RL) fine-tuning offers a promising path to bridge this gap, but existing approaches either train from scratch without fully leveraging pretrained priors, or fine-tune VLAs without achieving the sample efficiency and success rates that practical deployment demands. We present EXPO-FT, a system for stable, sample-efficient RL finetuning of pretrained VLA policies that closes this gap. Our system solves a suite of challenging manipulation tasks, including routing string lights and inserting the plug to light it up, striking a pool ball into a pocket, and inserting a flower into a wine bottle, each requiring combinations of high precision, dynamic actions, and robustness to varied initial states. Our system achieves perfect task performance (30/30 successes) across all evaluated tasks within an average of 19.1 minutes of online robot data, outperforming both prior RL-from-scratch and VLA finetuning approaches. We release an open-source codebase with the aim of facilitating broader adoption of RL finetuning of VLA models in robotics.

机器人强化学习视觉语言微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。