用结构语义地图和多模态大模型提升未知环境寻物成功率
Object Navigation with Structure-Semantic Reasoning-Based Multi-level Map and Multimodal Decision-Making LLM
- 构建环境属性地图,融合语义与扩散预测未见区域
- 在HM3D和MP3D上分别达到28.4%和26.3%的SPL,提升46.0%
- 适合长距离导航与零样本新目标场景的智能体研究
在未知开放环境中进行零样本物体导航(ZSON)时,因忽略高维隐含场景信息及长距离搜索任务,性能显著下降。为此,提出基于环境属性地图(EAM)与多模态大语言模型层次推理模块(MHR)的主动导航框架,以提升成功率与效率。EAM通过SBERT推理已观测环境,并利用扩散模型预测未观测区域,融合人类空间规律中的物体-房间关联与区域邻接关系。MHR受EAM启发,实现前沿探索决策,避免长距离场景下的迂回路径,提升路径效率。实验表明,EAM在MP3D数据集上达到64.5%的场景映射准确率;导航任务在HM3D与MP3D基准上分别获得28.4%与26.3%的SPL,相比基线方法绝对提升21.4%与46.0%。
原文摘要 · Abstract (English)
The zero-shot object navigation (ZSON) in unknown open-ended environments coupled with semantically novel target often suffers from the significant decline in performance due to the neglect of high-dimensional implicit scene information and the long-range target searching task. To address this, we proposed an active object navigation framework with Environmental Attributes Map (EAM) and MLLM Hierarchical Reasoning module (MHR) to improve its success rate and efficiency. EAM is constructed by reasoning observed environments with SBERT and predicting unobserved ones with Diffusion, utilizing human space regularities that underlie object-room correlations and area adjacencies. MHR is inspired by EAM to perform frontier exploration decision-making, avoiding the circuitous trajectories in long-range scenarios to improve path efficiency. Experimental results demonstrate that the EAM module achieves 64.5\% scene mapping accuracy on MP3D dataset, while the navigation task attains SPLs of 28.4\% and 26.3\% on HM3D and MP3D benchmarks respectively - representing absolute improvements of 21.4\% and 46.0\% over baseline methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。