arXiv:2603.16065cs.ROcs.AI2026-03被引 9

用视觉语言模型生成奖励,让机器人在线自适应修正操作错误。

Large Reward Models: Generalizable Online Robot Reward Generation with Vision-Language Models

  • 用大规模多源数据训练的视觉语言模型实时生成过程、完成度和时间对比奖励。
  • 仅30次强化学习迭代就显著提升初始模仿学习策略的成功率。
  • 零样本适配新任务,无需人工设计奖励函数,适合复杂长程操作场景。

强化学习在优化机器人操作策略方面潜力巨大,但其效果常受限于难以设计通用的奖励函数。本文提出一种框架,将基础视觉语言模型(VLM)转化为在线奖励生成器,用于持续优化机器人策略。我们基于先进VLM构建了一个稳健可扩展的奖励模型,该模型在包含真实机器人轨迹、人-物交互及多样仿真环境的大规模多源数据集上进行训练。与以往在事后评估整个轨迹的方法不同,本方法利用VLM根据当前视觉观测,生成包含过程、完成度和时间对比的多维度奖励信号。从模仿学习(IL)训练的基础策略出发,通过闭环方式使用VLM奖励引导模型修正次优行为。我们在需要序列执行与精确控制的长程操作基准上进行了评估,关键在于奖励模型在测试环境中完全零样本运行。实验结果表明,仅需30次强化学习迭代,该方法便显著提升了初始IL策略的成功率,展现出极高的样本效率。实证证明,由VLM生成的信号能提供可靠反馈以纠正执行错误,有效避免人工奖励工程,实现高效在线机器人学习优化。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) has shown great potential in refining robotic manipulation policies, yet its efficacy remains strongly bottlenecked by the difficulty of designing generalizable reward functions. In this paper, we propose a framework for online policy refinement by adapting foundation VLMs into online reward generators. We develop a robust, scalable reward model based on a state-of-the-art VLM, trained on a large-scale, multi-source dataset encompassing real-world robot trajectories, human-object interactions, and diverse simulated environments. Unlike prior approaches that evaluate entire trajectories post-hoc, our method leverages the VLM to formulate a multifaceted reward signal comprising process, completion, and temporal contrastive rewards based on current visual observations. Initializing with a base policy trained via Imitation Learning (IL), we employ these VLM rewards to guide the model to correct sub-optimal behaviors in a closed-loop manner. We evaluate our framework on challenging long-horizon manipulation benchmarks requiring sequential execution and precise control. Crucially, our reward model operates in a purely zero-shot manner within these test environments. Experimental results demonstrate that our method significantly improves the success rate of the initial IL policy within just 30 RL iterations, demonstrating remarkable sample efficiency. This empirical evidence highlights that VLM-generated signals can provide reliable feedback to resolve execution errors, effectively eliminating the need for manual reward engineering and facilitating efficient online refinement for robot learning.

机器人学习视觉语言模型强化学习在线优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。