arXiv:2506.06804cs.RO2025-06被引 5

融合激光雷达与视觉模型,快速构建高精度室内3D场景图。

IRS: Instance-Level 3D Scene Graphs via Room Prior Guided LiDAR-Camera Fusion

论文配图:IRS: Instance-Level 3D Scene Graphs via Room Prior Guided LiDAR-Camera Fusion
图 1 · 摘自论文原文
  • 用激光雷达获取房间级几何先验,加速场景理解
  • 多层级视觉基础模型提升语义识别准确率
  • 支持自然语言导航,适合机器人实际应用

室内场景理解是机器人领域的基础挑战,直接影响导航与操作等下游任务。传统方法依赖封闭集识别或回环检测,在开放世界中适应性差。随着视觉基础模型(VFMs)的发展,开放词汇识别与自然语言查询成为可能,为3D场景图构建带来新机遇。本文提出一种基于激光雷达-相机融合的实例级3D场景图构建框架。利用激光雷达的广视角与远距离感知能力,快速获取房间级几何先验。采用多层级视觉基础模型提升语义提取的准确性和一致性。在实例融合阶段,基于房间的分割实现并行处理,结合几何与语义线索显著提高融合精度与鲁棒性。相比最先进方法,本方案构建速度提升近一个数量级,同时保持高语义精度。在仿真与真实环境中的大量实验验证了方法有效性,并通过语言引导的语义导航任务展示了其实际应用潜力。

原文摘要 · Abstract (English)

Indoor scene understanding remains a fundamental challenge in robotics, with direct implications for downstream tasks such as navigation and manipulation. Traditional approaches often rely on closed-set recognition or loop closure, limiting their adaptability in open-world environments. With the advent of visual foundation models (VFMs), open-vocabulary recognition and natural language querying have become feasible, unlocking new possibilities for 3D scene graph construction. In this paper, we propose a robust and efficient framework for instance-level 3D scene graph construction via LiDAR-camera fusion. Leveraging LiDAR's wide field of view (FOV) and long-range sensing capabilities, we rapidly acquire room-level geometric priors. Multi-level VFMs are employed to improve the accuracy and consistency of semantic extraction. During instance fusion, room-based segmentation enables parallel processing, while the integration of geometric and semantic cues significantly enhances fusion accuracy and robustness. Compared to state-of-the-art methods, our approach achieves up to an order-of-magnitude improvement in construction speed while maintaining high semantic precision. Extensive experiments in both simulated and real-world environments validate the effectiveness of our approach. We further demonstrate its practical value through a language-guided semantic navigation task, highlighting its potential for real-world robotic applications.

3D场景图多模态融合机器人导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。