用在线强化学习提升视觉语言动作模型的实时适应能力
Improving Vision-Language-Action Model with Online Reinforcement Learning
- 交替使用强化学习与监督学习,兼顾探索性与稳定性
- 在仿真和真实机器人任务中均显著提升模型性能
- 适合希望改进大模型实时交互能力的研究者
近期研究通过监督微调大型视觉语言模型(VLM)并结合专家机器人数据,构建了视觉语言动作(VLA)模型。尽管性能强大,但如何在与环境交互过程中持续优化这些大模型仍是未解问题。本文探索利用强化学习(RL)进一步优化VLA模型。然而,直接应用在线RL面临训练不稳与计算负担过重等挑战。为此,我们提出iRe-VLA框架,通过迭代融合强化学习与监督学习,在保持稳定性的同时发挥强化学习的探索优势。在两个仿真基准和一个真实世界操作套件上的实验验证了该方法的有效性。
原文摘要 · Abstract (English)
Recent studies have successfully integrated large vision-language models (VLMs) into low-level robotic control by supervised fine-tuning (SFT) with expert robotic datasets, resulting in what we term vision-language-action (VLA) models. Although the VLA models are powerful, how to improve these large models during interaction with environments remains an open question. In this paper, we explore how to further improve these VLA models via Reinforcement Learning (RL), a commonly used fine-tuning technique for large models. However, we find that directly applying online RL to large VLA models presents significant challenges, including training instability that severely impacts the performance of large models, and computing burdens that exceed the capabilities of most local machines. To address these challenges, we propose iRe-VLA framework, which iterates between Reinforcement Learning and Supervised Learning to effectively improve VLA models, leveraging the exploratory benefits of RL while maintaining the stability of supervised learning. Experiments in two simulated benchmarks and a real-world manipulation suite validate the effectiveness of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。