用固定外置摄像头构建3D场景图先验,提升机器人探索效率
Fixed External Cameras as Common Prior Maps for Active 3D Scene Graph Generation

- 将外部摄像头视图为统一先验地图,融合进机器人感知流程
- 单个外部摄像头使初始物体召回率提升79%以上
- 适合需要高效建图的移动机器人场景
常见的先验信息如建筑信息模型(BIM)、平面图和遥感图像,可为自主机器人系统提供宝贵的几何与语义上下文。本文将固定外部RGB摄像头的观测视为通用先验地图(CPMs):在机器人开始运动前,提供环境的宽视野语义与几何先验。我们提出一种仅依赖RGB的主动增量式3D场景图(3DSG)生成框架,通过单一硬件无关的流程无缝融合机载相机与外部摄像头的观测。系统仅使用前馈3D重建模型处理所有摄像头输入,无需硬件改造。基于图结构的主动语义探索框架直接利用部分场景图引导机器人向高语义不确定性区域移动,逐步完善并优化先验。实验表明,仅用一个外部摄像头初始化场景图,初始物体召回率最高可提升79%,且先验提供的丰富上下文显著提升了后续主动探索的效率。
原文摘要 · Abstract (English)
Commonly available prior information, such as BIM models, floor plans, and remote sensing images, can provide valuable geometric and semantic context for autonomous robotic systems. In this paper, we treat observations from fixed external RGB cameras as Common Prior Maps (CPMs): wide-field views of the environment that initialize a semantic and geometric scene prior before any robot motion begins. We present an RGB-only framework for active, incremental 3D scene graph (3DSG) generation that seamlessly fuses observations from both onboard robot cameras and fixed external cameras within a single hardware-agnostic pipeline. By relying solely on RGB observations processed by a feed-forward 3D reconstruction model, the system treats all cameras - onboard or external - identically, requiring no hardware modifications. A graph-based active semantic exploration framework then directly leverages the partial scene graph to guide the robot toward regions of high semantic uncertainty, progressively completing and refining the prior. Experiments demonstrate that bootstrapping the scene graph with even a single external camera increases initial object recall by up to +79%, and that the richer context of the prior significantly improves the efficiency of subsequent active exploration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。