让机器人理解自由指令并精准导航,靠的是视觉语言的开放词汇映射。
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
- 用结构语义一致性约束融合几何与语义,实现3D实例鲁棒聚合。
- 通过大模型辅助解析指令,提升对目标物体的细粒度定位能力。
- 零样本下在扫描数据集上超越现有方法,适合智能导航研究者。
将自然语言指令与视觉观察进行语义对齐,是开放世界中具身智能体操作的基础。近期视觉-语言映射进展利用视觉语言模型(VLMs)实现了可泛化的语义表示,但这些方法在将自由形式的语言命令与具体场景实例对齐时仍存在不足,主要源于实例级语义一致性与指令理解能力的局限。本文提出 OpenMap,一种用于导航任务的零样本开放词汇视觉-语言地图,以实现准确的指令定位。为解决跨视角语义不一致问题,我们引入结构-语义共识约束,联合考虑全局几何结构与视觉-语言相似性,指导鲁棒的3D实例级聚合。为改善指令理解,我们提出基于大语言模型(LLM)的指令-实例对齐模块,通过结合空间上下文与丰富的目标描述,实现细粒度实例选择。我们在 ScanNet200 与 Matterport3D 上评估了 OpenMap,涵盖语义地图构建与指令-目标检索任务。实验结果表明,OpenMap 在零样本设置下优于当前最优基线,验证了该方法在连接自由语言与三维感知方面的有效性。
原文摘要 · Abstract (English)
Grounding natural language instructions to visual observations is fundamental for embodied agents operating in open-world environments. Recent advances in visual-language mapping have enabled generalizable semantic representations by leveraging vision-language models (VLMs). However, these methods often fall short in aligning free-form language commands with specific scene instances, due to limitations in both instance-level semantic consistency and instruction interpretation. We present OpenMap, a zero-shot open-vocabulary visual-language map designed for accurate instruction grounding in navigation tasks. To address semantic inconsistencies across views, we introduce a Structural-Semantic Consensus constraint that jointly considers global geometric structure and vision-language similarity to guide robust 3D instance-level aggregation. To improve instruction interpretation, we propose an LLM-assisted Instruction-to-Instance Grounding module that enables fine-grained instance selection by incorporating spatial context and expressive target descriptions. We evaluate OpenMap on ScanNet200 and Matterport3D, covering both semantic mapping and instruction-to-target retrieval tasks. Experimental results show that OpenMap outperforms state-of-the-art baselines in zero-shot settings, demonstrating the effectiveness of our method in bridging free-form language and 3D perception for embodied navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。