arXiv:2409.10849cs.ROcs.AI2024-09被引 5

让机器人听懂乱糟糟的人话,靠的是模拟人类的共情理解。

Pragmatic Embodied Spoken Instruction Following in Human-Robot Collaboration with Theory of Mind

  • 用视觉语言模型+心理推断,让机器人理解说话人意图
  • 在嘈杂环境中仍能接近人类准确率,比顶级模型还强
  • 适合做真实场景下人机协作的智能助手

语音指令在人机协作中无处不在。然而,在真实世界中,由于背景噪音或发音不准等因素,机器人理解人类口语面临挑战。人类可通过环境上下文对模糊语音指令进行语用推断并采取协助行为。本文提出一种受认知启发的神经符号模型SIFToM,利用基于模型的心理推断机制,结合视觉语言模型(VLM),使机器人在多种语音条件下实现语用性语音指令跟随。我们在虚拟环境(VirtualHome)和真实人机协作场景中进行了测试,并通过人工评估验证效果。结果表明,SIFToM显著提升了轻量级基线VLM(Gemini 2.5 Flash)的表现,超越当前最优的VLM(Gemini 2.5 Pro),在具有挑战性的语音指令任务中逼近人类水平准确率。

原文摘要 · Abstract (English)

Spoken language instructions are ubiquitous in agent collaboration. However, in real-world human-robot collaboration, following human spoken instructions can be challenging due to various speaker and environmental factors, such as background noise or mispronunciation. When faced with noisy auditory inputs, humans can leverage the collaborative context in the embodied environment to interpret noisy spoken instructions and take pragmatic assistive actions. In this paper, we present a cognitively inspired neurosymbolic model, Spoken Instruction Following through Theory of Mind (SIFToM), which leverages a Vision-Language Model with model-based mental inference to enable robots to pragmatically follow human instructions under diverse speech conditions. We test SIFToM in both simulated environments (VirtualHome) and real-world human-robot collaborative settings with human evaluations. Results show that SIFToM can significantly improve the performance of a lightweight base VLM (Gemini 2.5 Flash), outperforming state-of-the-art VLMs (Gemini 2.5 Pro) and approaching human-level accuracy on challenging spoken instruction following tasks.

人机协作语音理解心理建模视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。