arXiv:2505.07395cs.RO2025-05ICML被引 46

用强化学习提升机器人视觉语言模型的决策能力,尤其擅长处理低质量训练数据。

ReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement Learning

  • 通过预测密集回报,让模型理解任务中不同动作的质量差异。
  • 在CALVIN数据集上达到顶尖性能,少样本和分布外任务表现优异。
  • 适合需要鲁棒决策的机器人实际应用,尤其面对数据不齐的情况。

视觉-语言-动作(VLA)模型通过模仿学习在通用机器人决策任务中展现出巨大潜力,但训练数据质量参差不齐常制约其性能。相比之下,离线强化学习(RL)擅长从混合质量数据中学习稳健策略。本文提出一种新型端到端VLA模型ReinboT,融合强化学习中最大化累积奖励的原则。ReinboT通过预测密集回报,深入理解数据质量分布,捕捉操作任务中的细微差别,使机器人生成更稳健、面向未来收益最大化的决策动作。大量实验表明,ReinboT在包含混合质量数据的CALVIN数据集上达到当前最优性能,并在真实世界任务中展现出卓越的少样本学习与分布外泛化能力。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have shown great potential in general robotic decision-making tasks via imitation learning. However, the variable quality of training data often constrains the performance of these models. On the other hand, offline Reinforcement Learning (RL) excels at learning robust policy models from mixed-quality data. In this paper, we introduce Reinforced robot GPT (ReinboT), a novel end-to-end VLA model that integrates the RL principle of maximizing cumulative reward. ReinboT achieves a deeper understanding of the data quality distribution by predicting dense returns that capture the nuances of manipulation tasks. The dense return prediction capability enables the robot to generate more robust decision-making actions, oriented towards maximizing future benefits. Extensive experiments show that ReinboT achieves state-of-the-art performance on the CALVIN mixed-quality dataset and exhibits superior few-shot learning and out-of-distribution generalization capabilities in real-world tasks.

机器人强化学习视觉语言端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。