arXiv:2505.11214cs.RO2025-05被引 13

让机器人理解图片、视频等多模态指令,提升真实场景交互能力

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions

  • 构建支持多模态输入的VLA模型,可处理图像、视频等非语言指令
  • 在4类新任务上表现优异,性能接近纯语言指令模型
  • 适合需要灵活人机交互的智能机器人应用场景

视觉-语言-动作(VLA)模型在机器人领域日益重要。利用大规模网络数据训练的视觉-语言基础模型,VLA模型可通过单一端到端神经网络,直接从视觉观察和人类指令生成机器人动作。然而,现有VLA模型通常仅接受语言指令,限制了其在开放式人机交互中的应用。例如,用户可能希望机器人根据图片取物、读取白板文字或模仿视频中的行为,而不仅依赖语言描述。为此,我们提出OE-VLA,探索VLA模型在开放式多模态指令下的潜力。大量实验表明,OE-VLA在保持与传统语言输入模型相当性能的同时,在四类新增的开放任务中均取得出色表现。该方法显著拓展了VLA模型在日常场景中的应用范围,推动人机协作发展。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have recently become highly prominent in the field of robotics. Leveraging vision-language foundation models trained on large-scale internet data, the VLA model can generate robotic actions directly from visual observations and human instructions through a single end-to-end neural network. Despite their effectiveness, current VLA models usually accept only one form of human prompting, language instructions, which may constrain their applicability in open-ended human-robot interactions. For example, a user might expect the robot to retrieve an object shown in an image, follow an instruction written on the whiteboard, or imitate a behavior demonstrated in a video, rather than relying solely on language-based descriptions. To address this gap, we introduce OE-VLA, which explores the potential of VLA models for open-ended multimodal instructions. Extensive results demonstrate that our OE-VLA not only achieves comparable performance to traditional VLA models with linguistic input but also delivers impressive results across four additional categories of open-ended tasks. The proposed methodology could significantly expand the applications of VLA models across various everyday scenarios and facilitate human-robot interaction.

机器人多模态VLA人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。