实时构建带物体形状估计的3D场景图,提升环境重建精度。
Hydra++: Real-Time Hierarchical 3D Scene Graph Construction With Object-Level Shape Estimation

- 用学习型形状估计器替代粗略模板,实现对象级几何建模。
- 在真实校园环境中,显著提升物体与场景重建质量。
- 支持激光雷达-相机融合配置,适用于大规模户外场景。
3D场景图通过编码物体、场所及其关系,提供环境的层次化抽象。然而,现有系统对物体几何建模较粗糙,依赖部分点云或类别级CAD模板,限制了实例级细节。本文提出Hydra++,系统性研究如何将基于学习的物体形状估计器集成到层次化3D场景图管道中。Hydra++引入无类别形状估计和重投影掩码一致性检查,以剔除由部分观测或不精确分割导致的退化预测。其默认采用CRISP配置,实现在线场景图构建;评估了更慢的SAM3D等估算器作为模块化替代方案,展示泛化能力与延迟的权衡。此外,为应对室外环境中稀疏且嘈杂的深度测量问题,Hydra++支持混合激光雷达-相机配置,提升大范围场景重建质量。仿真与真实校园场景实验表明,Hydra++在物体级和场景级重建质量上均有显著提升。项目页面见https://hydra-plusplus.github.io/。
原文摘要 · Abstract (English)
3D scene graphs provide a hierarchical abstraction of environments by encoding spatial entities, such as objects and places, and their relationships. However, existing scene graph systems model object geometry coarsely, relying on partial point clouds or class-level CAD templates, which limits instance-specific shape detail. This paper presents Hydra++, a system-level investigation into how learning-based object shape estimators can be integrated into a hierarchical 3D scene graph pipeline. Hydra++ incorporates category-agnostic shape estimation and a reprojection-mask consistency check to reject degenerate predictions from partial observations or imprecise segmentation. In its default CRISP-based configuration, Hydra++ performs online scene graph construction; slower estimators such as SAM3D are evaluated as modular alternatives to demonstrate generalization-latency trade-offs. Furthermore, to address the challenges of sparse and noisy depth measurements in outdoor environments, Hydra++ supports a hybrid LiDAR-camera configuration for large-scale operation, improving scene-level reconstruction quality. Experiments in both simulation and real-world outdoor campus scenarios demonstrate that Hydra++ improves object- and scene-level reconstruction quality. Project page is available at https://hydra-plusplus.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。