arXiv:2507.04789cs.RO2025-07ICCV被引 4

无需训练,用视觉语言模型生成稳定奖励,提升机器人长时任务表现

Training-free Generation of Temporally Consistent Rewards from VLMs

  • 基于视觉语言模型的子目标追踪,动态更新任务完成度
  • 在两个机器人操作基准上达到最优性能,计算开销更低
  • 适合需要快速部署、无标注数据的具身智能场景

视觉语言模型(VLM)在具身任务中的目标分解和视觉理解方面取得了显著进展。然而,由于预训练数据中缺乏领域特定的机器人知识,且微调成本高,如何在不微调的前提下为机器人操作提供准确奖励仍具挑战。为此,我们提出T²-VLM,一种无需训练、时间一致的新型框架,通过追踪VLM生成的子目标状态变化来生成精准奖励。具体而言,每次交互前,先查询VLM以建立空间感知的子目标和初始完成度估计;随后采用贝叶斯追踪算法,利用子目标隐状态动态更新目标完成状态,生成结构化奖励用于强化学习(RL)代理。该方法增强了长周期决策能力,并提升了失败恢复性能。大量实验表明,T²-VLM在两个机器人操作基准上均达到当前最优表现,兼具更高的奖励准确性与更低的计算消耗。我们相信,该方法不仅推动了奖励生成技术的发展,也对具身人工智能具有广泛意义。

原文摘要 · Abstract (English)

Recent advances in vision-language models (VLMs) have significantly improved performance in embodied tasks such as goal decomposition and visual comprehension. However, providing accurate rewards for robotic manipulation without fine-tuning VLMs remains challenging due to the absence of domain-specific robotic knowledge in pre-trained datasets and high computational costs that hinder real-time applicability. To address this, we propose $\mathrm{T}^2$-VLM, a novel training-free, temporally consistent framework that generates accurate rewards through tracking the status changes in VLM-derived subgoals. Specifically, our method first queries the VLM to establish spatially aware subgoals and an initial completion estimate before each round of interaction. We then employ a Bayesian tracking algorithm to update the goal completion status dynamically, using subgoal hidden states to generate structured rewards for reinforcement learning (RL) agents. This approach enhances long-horizon decision-making and improves failure recovery capabilities with RL. Extensive experiments indicate that $\mathrm{T}^2$-VLM achieves state-of-the-art performance in two robot manipulation benchmarks, demonstrating superior reward accuracy with reduced computation consumption. We believe our approach not only advances reward generation techniques but also contributes to the broader field of embodied AI. Project website: https://t2-vlm.github.io/.

具身智能视觉语言模型强化学习机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。