arXiv:2605.13741cs.ROcs.CV2026-05被引 1

仅用RGB相机实现可扩展的3D场景图构建,突破传统依赖深度传感器的限制。

LEXI-SG: Monocular 3D Scene Graph Mapping with Room-Guided Feed-Forward Reconstruction

论文配图:LEXI-SG: Monocular 3D Scene Graph Mapping with Room-Guided Feed-Forward Reconstruction
图 1 · 摘自论文原文
  • 基于房间级分块与前馈重建,解决单目视觉尺度不一致问题
  • 在Habitat-Matterport和自采办公室数据集上实现精准轨迹估计与稠密重建
  • 支持开放词汇物体分割与追踪,适合机器人导航场景应用

场景图已成为机器人导航的标准表示,提供层次化的几何与语义场景理解。然而,现有方法多依赖深度相机或激光雷达。本文提出LEXI-SG,首个仅使用RGB相机输入的开放词汇3D场景图密集映射系统。该方法利用开放词汇基础模型的语义先验,将场景划分为房间,待每个房间完全观测后才进行前馈重建,从而实现无滑动窗口尺度不一致的可扩展稠密映射。我们设计基于房间的因子图公式,全局对齐房间重建,同时保持局部地图一致性,并自然引入语义场景图层级结构。每个房间内支持开放词汇物体分割与追踪。我们在Habitat-Matterport 3D及自采集的头戴式办公室序列中验证了LEXI-SG,在轨迹估计、稠密重建上优于现有前馈SLAM方法,开放词汇分割性能也具有竞争力。结果表明,仅凭单目RGB即可实现准确且可扩展的开放词汇3D场景图。项目主页与数据集链接:https://ori-drs.github.io/lexisg-web/

原文摘要 · Abstract (English)

Scene graphs are becoming a standard representation for robot navigation, providing hierarchical geometric and semantic scene understanding. However, most scene graph mapping methods rely on depth cameras or LiDAR sensors. In this work, we present LEXI-SG, the first dense monocular visual mapping system for open-vocabulary 3D scene graphs using only RGB camera input. Our approach exploits the semantic priors of open-vocabulary foundation models to partition the scene into rooms, deferring feed-forward reconstruction to when each room is fully observed -- enabling scalable dense mapping without sliding-window scale inconsistencies. We propose a room-based factor graph formulation to globally align room reconstructions while preserving local map consistency and naturally imposing the semantic scene graph hierarchy. Within each room, we further support open-vocabulary object segmentation and tracking. We validate LEXI-SG on indoor scenes from the Habitat-Matterport 3D and self-collected egocentric office sequences. We evaluate its performance against existing feed-forward SLAM methods, as well as established scene graphs baselines. We demonstrate improved trajectory estimation and dense reconstruction, as well as, competitive performance in open-vocabulary segmentation. LEXI-SG shows that accurate, scalable, open-vocabulary 3D scene graphs can be achieved from monocular RGB alone. Our project page and office sequences are available here: https://ori-drs.github.io/lexisg-web/.

3D场景图单目视觉开放词汇机器人导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。