用360°无人机影像重建3D场景并精准识别室内资产
Indoor Asset Detection in Large Scale 360° Drone-Captured Imagery via 3D Gaussian Splatting
- 构建3D物体代码本,融合语义与空间信息实现多视角掩码关联
- 在两个大型室内场景上实现65%的F1分数提升和11%的mAP增益
- 适合需要高精度室内资产检测的智能建筑与数字孪生应用
我们提出一种在360°无人机采集影像重建的3D Gaussian Splatting(3DGS)场景中,对目标室内资产进行逐对象检测与分割的方法。引入3D物体代码本,联合利用对应高斯原语的掩码语义与空间信息,指导多视角掩码关联与室内资产检测。通过集成2D目标检测与分割模型,并结合语义与空间约束的合并流程,将多视角掩码聚合为连贯的3D物体实例。在两个大型室内场景上的实验表明,该方法实现了可靠的多视角掩码一致性,相比最先进基线提升65%的F1分数;同时在物体级3D室内资产检测任务中,达到11%的mAP增益。
原文摘要 · Abstract (English)
We present an approach for object-level detection and segmentation of target indoor assets in 3D Gaussian Splatting (3DGS) scenes, reconstructed from 360° drone-captured imagery. We introduce a 3D object codebook that jointly leverages mask semantics and spatial information of their corresponding Gaussian primitives to guide multi-view mask association and indoor asset detection. By integrating 2D object detection and segmentation models with semantically and spatially constrained merging procedures, our method aggregates masks from multiple views into coherent 3D object instances. Experiments on two large indoor scenes demonstrate reliable multi-view mask consistency, improving F1 score by 65% over state-of-the-art baselines, and accurate object-level 3D indoor asset detection, achieving an 11% mAP gain over baseline methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。