用多模态信息构建可导航的语义地图,让机器人听懂复杂指令
Multimodal Spatial Language Maps for Robot Navigation and Manipulation

- 融合视觉、音频与语言特征,生成带空间位置的多模态地图
- 在模糊场景中召回率提升50%,支持零样本跨设备导航
- 适合需要听懂口语指令的移动机器人和操作机械臂
将语言与导航智能体的感知进行对齐,可利用预训练多模态基础模型将感知与物体或事件描述匹配。然而,现有方法在环境建图方面仍存在脱节,缺乏几何地图的空间精度,或忽略视觉之外的模态信息。为此,我们提出多模态空间语言地图,一种融合预训练多模态特征与环境三维重建的空间地图表示。通过标准探索过程自主构建该地图。我们提出了两种实例:视觉-语言地图(VLMaps)及其扩展——音频-视觉-语言地图(AVLMaps),后者通过加入音频信息实现。结合大语言模型(LLMs),VLMaps可直接将自然语言命令(如“沙发与电视之间”)转化为开放词汇空间目标并定位到地图中;且可跨不同机器人形态共享,按需生成定制障碍物地图。在此基础上,AVLMaps通过融合预训练多模态基础模型的特征,建立统一的3D空间表示,整合音频、视觉与语言线索,使机器人能将多模态目标查询(如文本、图像或音频片段)精准定位至空间位置。此外,多样感官输入显著提升在模糊环境中的目标消歧能力。仿真与真实世界实验表明,我们的多模态空间语言地图支持零样本空间与多模态目标导航,在模糊场景中召回率提升50%。该能力适用于移动机器人和桌面操作机械臂,支持由视觉、音频及空间线索引导的导航与交互。
原文摘要 · Abstract (English)
Grounding language to a navigating agent's observations can leverage pretrained multimodal foundation models to match perceptions to object or event descriptions. However, previous approaches remain disconnected from environment mapping, lack the spatial precision of geometric maps, or neglect additional modality information beyond vision. To address this, we propose multimodal spatial language maps as a spatial map representation that fuses pretrained multimodal features with a 3D reconstruction of the environment. We build these maps autonomously using standard exploration. We present two instances of our maps, which are visual-language maps (VLMaps) and their extension to audio-visual-language maps (AVLMaps) obtained by adding audio information. When combined with large language models (LLMs), VLMaps can (i) translate natural language commands into open-vocabulary spatial goals (e.g., "in between the sofa and TV") directly localized in the map, and (ii) be shared across different robot embodiments to generate tailored obstacle maps on demand. Building upon the capabilities above, AVLMaps extend VLMaps by introducing a unified 3D spatial representation integrating audio, visual, and language cues through the fusion of features from pretrained multimodal foundation models. This enables robots to ground multimodal goal queries (e.g., text, images, or audio snippets) to spatial locations for navigation. Additionally, the incorporation of diverse sensory inputs significantly enhances goal disambiguation in ambiguous environments. Experiments in simulation and real-world settings demonstrate that our multimodal spatial language maps enable zero-shot spatial and multimodal goal navigation and improve recall by 50% in ambiguous scenarios. These capabilities extend to mobile robots and tabletop manipulators, supporting navigation and interaction guided by visual, audio, and spatial cues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。