arXiv:2512.01009cs.ROcs.CV2025-12被引 2

用前沿物体地图提升机器人找物效率

FOM-Nav: Frontier-Object Maps for Object Goal Navigation

  • 构建在线更新的前沿-物体地图,融合空间边界与物体细节
  • 视觉语言模型实现语义理解,提升目标预测准确率
  • 在真实机器人上表现优异,适合复杂环境导航任务

本文解决机器人在未知环境中寻找目标物体的问题。现有基于隐式记忆的方法难以长期保留信息和规划,而显式地图方法缺乏丰富语义。为此,我们提出FOM-Nav框架,通过在线构建的前沿-物体地图,联合编码空间前沿与细粒度物体信息。结合视觉语言模型实现多模态场景理解与高层目标预测,并由底层规划器生成高效轨迹。为训练该模型,我们从真实扫描环境自动生成大规模导航数据集。大量实验验证了模型设计与数据集的有效性。FOM-Nav在MP3D和HM3D基准上达到领先性能,尤其在导航效率指标SPL上表现突出,并在真实机器人上取得良好结果。

原文摘要 · Abstract (English)

This paper addresses the Object Goal Navigation problem, where a robot must efficiently find a target object in an unknown environment. Existing implicit memory-based methods struggle with long-term memory retention and planning, while explicit map-based approaches lack rich semantic information. To address these challenges, we propose FOM-Nav, a modular framework that enhances exploration efficiency through Frontier-Object Maps and vision-language models. Our Frontier-Object Maps are built online and jointly encode spatial frontiers and fine-grained object information. Using this representation, a vision-language model performs multimodal scene understanding and high-level goal prediction, which is executed by a low-level planner for efficient trajectory generation. To train FOM-Nav, we automatically construct large-scale navigation datasets from real-world scanned environments. Extensive experiments validate the effectiveness of our model design and constructed dataset. FOM-Nav achieves state-of-the-art performance on the MP3D and HM3D benchmarks, particularly in navigation efficiency metric SPL, and yields promising results on a real robot.

机器人导航视觉语言模型地图构建目标导向

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。