通过跨模态推理融合语义与视觉信息,提升机器人在弱语义环境下的导航能力。
SONAR: Semantic-Object Navigation with Aggregated Reasoning through a Cross-Modal Inference Paradigm
- 融合语义地图与视觉语言模型,实现多模态联合推理
- 在MP3D数据集上达成38.4%成功率和17.7%SPL
- 适合需强泛化与场景适应性的智能导航任务
理解人类指令并在未知环境中完成视觉语言导航对机器人至关重要。现有模块化方法依赖训练数据质量,泛化能力差;基于视觉语言模型的方法虽泛化性强,但在语义线索弱时表现不佳。本文提出SONAR,一种通过跨模态推理实现聚合推理的方法。该方法结合基于语义地图的目标预测模块与基于视觉语言模型的价值图模块,提升了在不同语义强度环境下的鲁棒导航能力,并有效平衡泛化性与场景适应性。针对目标定位,提出将多尺度语义地图与置信度图融合的策略,以减少目标误检。在Gazebo仿真器中,以最具挑战性的Matterport 3D(MP3D)数据集为基准进行评估。实验结果表明,SONAR在该数据集上达到38.4%的成功率和17.7%的SPL。
原文摘要 · Abstract (English)
Understanding human instructions and accomplishing Vision-Language Navigation tasks in unknown environments is essential for robots. However, existing modular approaches heavily rely on the quality of training data and often exhibit poor generalization. Vision-Language Model based methods, while demonstrating strong generalization capabilities, tend to perform unsatisfactorily when semantic cues are weak. To address these issues, this paper proposes SONAR, an aggregated reasoning approach through a cross modal paradigm. The proposed method integrates a semantic map based target prediction module with a Vision-Language Model based value map module, enabling more robust navigation in unknown environments with varying levels of semantic cues, and effectively balancing generalization ability with scene adaptability. In terms of target localization, we propose a strategy that integrates multi-scale semantic maps with confidence maps, aiming to mitigate false detections of target objects. We conducted an evaluation of the SONAR within the Gazebo simulator, leveraging the most challenging Matterport 3D (MP3D) dataset as the experimental benchmark. Experimental results demonstrate that SONAR achieves a success rate of 38.4% and an SPL of 17.7%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。