arXiv:2411.04549cs.ROcs.AI2024-11ICLR被引 110

利用视觉语言模型的常识知识,零样本预测机器人任务进展。

Vision Language Models are In-Context Value Learners

  • 将价值预测转为打乱帧的时序排序任务,激发模型语义与时间理解能力。
  • 无需训练,在300多个真实任务上实现零样本/少样本有效价值预测。
  • 支持多模态示例学习,适用于机器人、人形动作等跨任务场景。

从视觉轨迹预测时间进展对可学习、可适应的智能机器人至关重要。然而,跨任务和领域学习此类进度估计器(即时间价值函数)需要大量多样数据以及可扩展、可泛化的方法。为此,我们提出生成式价值学习(GVL),一种利用视觉语言模型(VLMs)中嵌入的世界知识来预测任务进展的通用价值函数估计器。直接让VLM对视频序列预测价值表现不佳,因连续帧存在强时序相关性。GVL将价值估计转化为对打乱视频帧的时序排序问题;这一看似更难的任务促使VLM更充分地利用其内在的语义与时间基础能力,基于感知任务进展区分帧,从而显著提升价值预测效果。无需任何机器人或任务特定训练,GVL可在上下文内实现零样本与少样本预测,覆盖超过300种不同现实任务,涵盖复杂双臂操作任务。此外,我们证明了GVL可通过异构任务和实体的示例实现灵活的多模态上下文学习,例如人类视频。GVL的通用性使其能支持多种下游应用,包括数据集筛选、成功检测和优势加权回归——均无需模型训练或微调。

原文摘要 · Abstract (English)

Predicting temporal progress from visual trajectories is important for intelligent robots that can learn, adapt, and improve. However, learning such progress estimator, or temporal value function, across different tasks and domains requires both a large amount of diverse data and methods which can scale and generalize. To address these challenges, we present Generative Value Learning (\GVL), a universal value function estimator that leverages the world knowledge embedded in vision-language models (VLMs) to predict task progress. Naively asking a VLM to predict values for a video sequence performs poorly due to the strong temporal correlation between successive frames. Instead, GVL poses value estimation as a temporal ordering problem over shuffled video frames; this seemingly more challenging task encourages VLMs to more fully exploit their underlying semantic and temporal grounding capabilities to differentiate frames based on their perceived task progress, consequently producing significantly better value predictions. Without any robot or task specific training, GVL can in-context zero-shot and few-shot predict effective values for more than 300 distinct real-world tasks across diverse robot platforms, including challenging bimanual manipulation tasks. Furthermore, we demonstrate that GVL permits flexible multi-modal in-context learning via examples from heterogeneous tasks and embodiments, such as human videos. The generality of GVL enables various downstream applications pertinent to visuomotor policy learning, including dataset filtering, success detection, and advantage-weighted regression -- all without any model training or finetuning.

价值学习视觉语言模型机器人零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。