仅用RGB相机实现室内机器人主动构建3D场景图,无需深度传感器。
RGB-only Active 3D Scene Graph Generation for Indoor Mobile Robots

- 基于共享结构化表示统一感知与规划,支持多视角融合。
- 在Replica上达到与依赖真实深度的基线相同F1分数。
- 语义驱动视角选择可提升探测物体数超2倍,适合资源受限场景。
当前3D场景图生成方法依赖专用深度传感器(如LiDAR或RGB-D相机)进行度量三维重建,限制了部署范围,难以适用于仅配备RGB相机的平台(如固定外部摄像头)。现有流程通常基于被动采集的观测轨迹,无法根据已构建的场景表示主动选择视点,因而未能有效利用图中蕴含的语义与空间信息。本文提出一种全视觉框架,仅使用RGB输入即可实现主动、增量式构建3D场景图,同时解决上述问题。该方法围绕共享结构化表示统一感知与规划,融合对象语义、3D几何、关系上下文及多视角信息。由于硬件无关且仅依赖RGB观测,系统可整合机器人机载相机与固定外部相机的数据于同一表示中。在Replica数据集上的实验表明,该纯RGB流水线在F1分数上与使用真实深度的基线相当;在ReplicaCAD上的主动探索实验显示,在相同探索预算下,语义驱动的视角选择比基于几何前沿的基线多检测超过两倍的物体;此外,外部相机提供的互补视角可有效启动场景图构建并提升上下文理解,且无需额外探索成本。
原文摘要 · Abstract (English)
Current approaches to 3D scene graph generation rely on dedicated depth sensors, such as LiDAR or RGB-D cameras, for metric 3D reconstruction. This limits deployment to specialized robotic platforms and excludes settings where only RGB cameras are available, such as fixed external infrastructure. Existing pipelines also typically operate on passively collected observation trajectories, rather than selecting viewpoints based on the partially built scene representation, and therefore fail to effectively exploit the semantic and spatial information encoded within the graph during exploration. This paper presents a fully visual framework for the active, incremental construction of 3D scene graphs from RGB input only, addressing both limitations. The proposed approach unifies perception and planning around a shared structured representation that captures object semantics, 3D geometry, relational context, and information from multiple viewpoints. Because the framework is hardware-agnostic and relies only on RGB observations, it can incorporate inputs from both onboard robot cameras and fixed external cameras within the same representation. Experiments on the Replica dataset show that the RGB-only pipeline achieves F1-score parity with baselines using ground-truth depth. Active exploration experiments on ReplicaCAD further show that semantic-driven viewpoint selection detects more than twice as many objects as a geometric frontier-based baseline under the same exploration budget. Finally, the external-camera setting demonstrates that complementary RGB views can effectively bootstrap the scene graph and improve contextual understanding at no additional exploration cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。