用真实机器人操作数据生成视觉问答数据集,提升模型空间与交互推理能力。
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
- 基于机器人轨迹中的位姿、夹持器开合等多模态数据,分阶段生成问答任务。
- 构建包含68.5万问的Robo2VLM-1数据集,覆盖3396项任务和463个场景。
- 适合研究机器人感知、视觉语言模型评估与具身智能的学者使用。
视觉语言模型(VLMs)通过互联网规模的图文语料获得现实世界知识与通用推理能力,可增强机器人系统的场景理解与任务规划,并辅助基于机器人轨迹训练的视觉运动策略。本文探索反向范式——利用丰富、真实、多模态的机器人轨迹数据来增强和评估VLM。提出Robo2VLM框架,用于生成视觉问答(VQA)数据集。给定人类遥控的机器人轨迹,Robo2VLM从非视觉、非描述性传感模态(如末端执行器位姿、夹持器开合、力觉)中提取真实答案,将轨迹分割为一系列操作阶段。在每个阶段,结合场景与交互理解,识别机器人、任务目标及目标物体的三维属性。基于这些属性,采用空间、目标条件与交互推理模板生成具有文本选项的图像问答问题。我们构建了Robo2VLM-1,一个大规模真实场景数据集,包含684,710个问题,覆盖463个不同场景与3,396项机器人操作任务,源自176,000条真实机器人轨迹。实验表明,Robo2VLM-1可用于评估并提升VLM在空间与交互推理方面的能力。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) acquire real-world knowledge and general reasoning ability through Internet-scale image-text corpora. They can augment robotic systems with scene understanding and task planning, and assist visuomotor policies that are trained on robot trajectory data. We explore the reverse paradigm - using rich, real, multi-modal robot trajectory data to enhance and evaluate VLMs. In this paper, we present Robo2VLM, a Visual Question Answering (VQA) dataset generation framework for VLMs. Given a human tele-operated robot trajectory, Robo2VLM derives ground-truth from non-visual and non-descriptive sensory modalities, such as end-effector pose, gripper aperture, and force sensing. Based on these modalities, it segments the robot trajectory into a sequence of manipulation phases. At each phase, Robo2VLM uses scene and interaction understanding to identify 3D properties of the robot, task goal, and the target object. The properties are used to generate representative VQA queries - images with textural multiple-choice questions - based on spatial, goal-conditioned, and interaction reasoning question templates. We curate Robo2VLM-1, a large-scale in-the-wild dataset with 684,710 questions covering 463 distinct scenes and 3,396 robotic manipulation tasks from 176k real robot trajectories. Results suggest that Robo2VLM-1 can benchmark and improve VLM capabilities in spatial and interaction reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。