arXiv:2512.20876cs.RO2025-12被引 1

用机器人运动数据增强视觉语言模型,提升任务描述与子任务分割能力

Proprioception Enhances Vision Language Model in Generating Captions and Subtask Segmentations for Robot Task

  • 融合图像描述与机器人关节/末端状态轨迹生成场景化字幕
  • 通过文本嵌入相似度实现子任务自动分割,提升任务理解精度
  • 适用于需语言-运动对齐的机器人模仿学习场景

从机器人未来发展角度看,验证仅基于离线图像和语言数据训练的基础模型是否能理解机器人运动至关重要。由于视觉语言模型(VLMs)的训练数据未包含机器人低层运动信息,包含轨迹信息的视频理解仍具挑战性。本研究通过带低层机器人运动信息的视频字幕任务,评估VLM的两项能力:(1)机器人任务的自动字幕生成;(2)一系列任务的子任务分割。这两项能力可提升机器人模仿学习效率,实现语言与运动的关联,并作为基础模型性能的衡量指标。所提方法利用图像字幕与机器人任务轨迹数据生成多个“场景”字幕,再通过总结生成完整任务字幕。同时,通过比较图像字幕文本嵌入的相似性完成子任务分割。在两种字幕任务中,均将机器人的运动数据(关节状态与末端执行器状态)作为输入以提升VLM表现。通过模拟器实验验证了该方法的有效性。

原文摘要 · Abstract (English)

From the perspective of future developments in robotics, it is crucial to verify whether foundation models trained exclusively on offline data, such as images and language, can understand the robot motion. In particular, since Vision Language Models (VLMs) do not include low-level motion information from robots in their training datasets, video understanding including trajectory information remains a significant challenge. In this study, we assess two capabilities of VLMs through a video captioning task with low-level robot motion information: (1) automatic captioning of robot tasks and (2) segmentation of a series of tasks. Both capabilities are expected to enhance the efficiency of robot imitation learning by linking language and motion and serve as a measure of the foundation model's performance. The proposed method generates multiple "scene" captions using image captions and trajectory data from robot tasks. The full task caption is then generated by summarizing these individual captions. Additionally, the method performs subtask segmentation by comparing the similarity between text embeddings of image captions. In both captioning tasks, the proposed method aims to improve performance by providing the robot's motion data - joint and end-effector states - as input to the VLM. Simulator experiments were conducted to validate the effectiveness of the proposed method.

视觉语言模型机器人运动任务分割模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。