arXiv:2506.15402cs.ROcs.AI2025-06被引 3

多视角全景相机实现复杂户外环境下的精准物体级地图构建。

MCOO-SLAM: A Multi-Camera Omnidirectional Object SLAM System

  • 用多摄像头全景配置融合点特征与语义物体地标。
  • 跨视角关联更鲁棒,支持遮挡和姿态变化下的稳定建图。
  • 适合需要语义理解的机器人导航与场景推理任务。

物体级SLAM提供结构化且语义丰富的环境表示,更适合高层机器人任务。然而,现有方法多依赖RGB-D传感器或单目视觉,受限于视场狭窄、易受遮挡及深度感知不足,尤其在大规模或室外场景中表现不佳,常只能从有限视角观测物体,导致建模不准和数据关联不可靠。本文提出MCOO-SLAM,一种新型多相机全景物体级SLAM系统,充分利用环视相机配置,在复杂室外场景中实现鲁棒、一致且语义丰富的建图。方法融合点特征与带开放词汇语义增强的物体级地物,采用语义-几何-时间融合策略实现跨视角鲁棒物体关联,提升一致性与建模精度;设计全景环路闭合模块,利用场景级描述符实现视角不变的位置识别。此外,构建的地图被抽象为层次化3D场景图,支持下游推理任务。大量真实世界实验表明,MCOO-SLAM在定位精度与可扩展物体级建图方面表现优异,对遮挡、姿态变化和环境复杂性具有更强鲁棒性。

原文摘要 · Abstract (English)

Object-level SLAM offers structured and semantically meaningful environment representations, making it more interpretable and suitable for high-level robotic tasks. However, most existing approaches rely on RGB-D sensors or monocular views, which suffer from narrow fields of view, occlusion sensitivity, and limited depth perception-especially in large-scale or outdoor environments. These limitations often restrict the system to observing only partial views of objects from limited perspectives, leading to inaccurate object modeling and unreliable data association. In this work, we propose MCOO-SLAM, a novel Multi-Camera Omnidirectional Object SLAM system that fully leverages surround-view camera configurations to achieve robust, consistent, and semantically enriched mapping in complex outdoor scenarios. Our approach integrates point features and object-level landmarks enhanced with open-vocabulary semantics. A semantic-geometric-temporal fusion strategy is introduced for robust object association across multiple views, leading to improved consistency and accurate object modeling, and an omnidirectional loop closure module is designed to enable viewpoint-invariant place recognition using scene-level descriptors. Furthermore, the constructed map is abstracted into a hierarchical 3D scene graph to support downstream reasoning tasks. Extensive experiments in real-world demonstrate that MCOO-SLAM achieves accurate localization and scalable object-level mapping with improved robustness to occlusion, pose variation, and environmental complexity.

SLAM物体级建图多相机语义感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。