arXiv:2502.13508cs.RO2025-02ICLR被引 61

让机器人直接听懂口语指令并执行定制化操作

VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation

  • 将语音识别嵌入机器人策略模型,实现端到端语音理解
  • 在新数据集上训练后,能准确响应多样化口语指令
  • 支持个性化语音知识检索,适合人机交互场景

视觉-语言-动作模型(VLAs)在机器人操作中因端到端设计和出色表现日益流行。然而,现有VLAs严重依赖仅支持文本指令的视觉-语言模型,忽视了更自然的语音交互方式。传统语音集成方法通常需独立语音识别系统,增加复杂性并引入误差传播。此外,转写过程会丢失语音中的非语义信息(如音色),而这些信息对完成定制任务可能至关重要。为此,我们提出VLAS,一种新型端到端视觉-语言-动作模型,将语音识别直接融入机器人策略模型。VLAS通过内部语音-文本对齐理解口语指令,并生成相应动作。我们还构建了两个新数据集SQA和CSI,支持三阶段微调流程,使模型具备跨文本、图像、语音与机器人动作的多模态交互能力。进一步地,设计了基于语音检索增强生成(RAG)的范式,使模型能够有效处理需要个体特定知识的任务。大量实验表明,VLAS可成功完成多种口语指令下的机器人操作任务,提供无缝且个性化的交互体验。

原文摘要 · Abstract (English)

Vision-language-action models (VLAs) have become increasingly popular in robot manipulation for their end-to-end design and remarkable performance. However, existing VLAs rely heavily on vision-language models (VLMs) that only support text-based instructions, neglecting the more natural speech modality for human-robot interaction. Traditional speech integration methods usually involves a separate speech recognition system, which complicates the model and introduces error propagation. Moreover, the transcription procedure would lose non-semantic information in the raw speech, such as voiceprint, which may be crucial for robots to successfully complete customized tasks. To overcome above challenges, we propose VLAS, a novel end-to-end VLA that integrates speech recognition directly into the robot policy model. VLAS allows the robot to understand spoken commands through inner speech-text alignment and produces corresponding actions to fulfill the task. We also present two new datasets, SQA and CSI, to support a three-stage tuning process for speech instructions, which empowers VLAS with the ability of multimodal interaction across text, image, speech, and robot actions. Taking a step further, a voice retrieval-augmented generation (RAG) paradigm is designed to enable our model to effectively handle tasks that require individual-specific knowledge. Our extensive experiments show that VLAS can effectively accomplish robot manipulation tasks with diverse speech commands, offering a seamless and customized interaction experience.

机器人操作语音指令多模态个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。