arXiv:2505.17016cs.LGcs.AI2025-05被引 115

用极少数据让视觉语言动作模型学会新任务,成功率最高达97.5%

Interactive Post-Training for Vision-Language-Action Models

  • 基于强化学习和稀疏成功奖励,实现交互式后训练
  • 仅需一次示范,4%成功率模型提升至97%成功
  • 适用于多种模型,泛化性强且对初始状态不敏感

我们提出RIPT-VLA,一种基于强化学习的简单可扩展交互式后训练范式,仅使用稀疏二元成功奖励微调预训练视觉-语言-动作(VLA)模型。现有VLA训练依赖大量离线专家示范数据和监督模仿,难以在低数据场景下适应新任务与环境。RIPT-VLA通过动态回溯采样和留一法优势估计的稳定策略优化算法,解决该问题。该方法适用于多种VLA模型,使轻量级QueST模型性能提升21.2%,7B OpenVLA-OFT模型达到前所未有的97.5%成功率。其计算与数据效率高:仅需一次示范,即可将原本失败率高达96%的SFT模型(原成功率4%)在15次迭代内提升至97%成功率。此外,所学策略具备跨任务与场景的泛化能力,对初始状态上下文具有鲁棒性。结果表明,RIPT-VLA是一种通过极少量监督实现高效后训练的有效范式。

原文摘要 · Abstract (English)

We introduce RIPT-VLA, a simple and scalable reinforcement-learning-based interactive post-training paradigm that fine-tunes pretrained Vision-Language-Action (VLA) models using only sparse binary success rewards. Existing VLA training pipelines rely heavily on offline expert demonstration data and supervised imitation, limiting their ability to adapt to new tasks and environments under low-data regimes. RIPT-VLA addresses this by enabling interactive post-training with a stable policy optimization algorithm based on dynamic rollout sampling and leave-one-out advantage estimation. RIPT-VLA has the following characteristics. First, it applies to various VLA models, resulting in an improvement on the lightweight QueST model by 21.2%, and the 7B OpenVLA-OFT model to an unprecedented 97.5% success rate. Second, it is computationally efficient and data-efficient: with only one demonstration, RIPT-VLA enables an unworkable SFT model (4%) to succeed with a 97% success rate within 15 iterations. Furthermore, we demonstrate that the policy learned by RIPT-VLA generalizes across different tasks and scenarios and is robust to the initial state context. These results highlight RIPT-VLA as a practical and effective paradigm for post-training VLA models through minimal supervision.

视觉语言动作强化学习后训练低数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。