仅用摄像头和惯性传感器,实时构建室内场景语义图谱。
Mono-Hydra++: Real-Time Monocular Scene Graph Construction with Multi-Task Learning for 3D Indoor Mapping

- 融合多任务学习与视觉惯性里程计,实现单目实时语义建图。
- 在ScanNet上轨迹误差比最强基线低1.6%,7-Scenes提升29.8%。
- 可在Jetson Orin NX上以25.53帧/秒运行,适合轻量机器人部署。
自主敏捷机器人不仅需要度量几何信息,还需理解物体、房间、位置及空间关系,以支持搜索、检查、探索和人机交互。传统度量地图虽能支撑定位与避障,但缺乏语义与关系结构。3D场景图通过连接几何与物体级、房间级理解填补此空白。然而,空中与轻量机器人受载荷、功耗和计算限制,难以使用RGB-D相机或激光雷达。本文提出Mono-Hydra++,一种基于单目RGB加IMU的实时室内度量语义建图与分层3D场景图构建系统。该系统结合M2H-MX(基于DINOv3的多任务模型)进行深度与语义估计,采用深度特征视觉惯性里程计前端,利用VIO生成的位姿图施加稀疏预测深度约束,通过语义掩码处理动态区域,并在体积融合前进行姿态感知的时间对齐。在Go-SLAM ScanNet评估子集上,相比最强的RGB-D基线,平均轨迹误差降低1.6%;在标定的7-Scenes数据集上,平均ATE提升29.8%。进一步在真实ITC建筑环境中部署RealSense RGB+IMU,验证其嵌入可行性:将ONNX/TensorRT FP16 M2H-MX-L感知模型在Jetson Orin NX 16GB上以25.53 FPS运行。结果表明,Mono-Hydra++可在无主动深度传感器条件下,为资源受限平台提供实时度量语义建图与场景图构建。
原文摘要 · Abstract (English)
Autonomous agile robots need more than metric geometry: they must understand objects, rooms, places, and spatial relations for search, inspection, exploration, and human robot interaction. Conventional metric maps support localization and collision avoidance, but do not provide this semantic and relational structure. 3D scene graphs address this gap by connecting geometry with object level and room level understanding. Building such representations on agile platforms remains difficult because aerial and lightweight robots operate under strict payload, power, and compute limits, making RGB-D cameras and LiDAR sensors impractical for many onboard settings. We present Mono-Hydra++, a real time monocular RGB plus IMU pipeline for indoor metric semantic mapping and hierarchical 3D scene graph construction. The system combines M2H-MX, a DINOv3 based multi-task model for depth and semantics, with a deep feature visual inertial odometry front end, sparse predicted depth constraints in the VIO derived pose graph, semantic masking for dynamic regions, and pose aware temporal alignment before volumetric fusion in the Mono-Hydra backend. On the Go-SLAM ScanNet evaluation subset, Mono-Hydra++ achieves 1.6% lower average trajectory error than the strongest RGB-D baseline in our comparison, while using only monocular RGB plus IMU input. On calibrated 7-Scenes, it improves average ATE by 29.8% over the strongest competing calibrated baseline. We further validate Mono-Hydra++ in a real ITC building deployment using RealSense RGB plus IMU and demonstrate embedded feasibility by deploying the ONNX/TensorRT FP16 M2H-MX-L perception model at 25.53 FPS on a Jetson Orin NX 16GB. These results show that Mono-Hydra++ can provide real time metric semantic mapping and scene graph construction for resource constrained robotic platforms without relying on active depth sensors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。