用卫星图和地图数据辅助车载摄像头补全3D场景,提升重建精度。
Geospatial-Prior Guidance for 3D Semantic Scene Completion

- 融合卫星影像与开放街道地图作为先验,指导3D场景补全。
- 在SemanticKITTI和SSCBench-KITTI-360上几何与语义完成度均提升。
- 特别适用于大范围静态结构如道路和建筑的补全,适合自动驾驶场景。
从车载图像中推断完整的3D几何与语义仍具挑战性,因遮挡和视野受限导致大量区域信息不足。尽管卫星影像提供大范围上下文,但仅依赖外观线索结构引导有限,且可能因时空差异不可靠。我们提出GeoScene框架,联合使用卫星影像和结构化OpenStreetMap作为软先验,用于3D语义场景补全。GeoScene学习车载观测与地理空间引导的体素级可靠性权重,并据此控制可观测与未观测区域的特征优化。该设计在保留局部视觉证据的同时,利用车载视野外的大规模道路与建筑结构。在SemanticKITTI和SSCBench-KITTI-360上的实验表明,GeoScene在地理空间先验辅助设定下持续提升几何与语义补全效果,对大尺度静态及地理结构化类别改善最为显著。
原文摘要 · Abstract (English)
Inferring complete 3D geometry and semantics from onboard images remains challenging because occlusions and restricted fields of view leave large scene regions underconstrained. Although satellite imagery provides wide-area context, appearance cues alone offer limited structural guidance and may be unreliable because of spatial or temporal discrepancies. We present GeoScene, a geospatially guided framework that jointly uses satellite imagery and structured OpenStreetMap cues as soft priors for 3D semantic scene completion. GeoScene learns complementary voxel-wise reliability weights for onboard observations and geospatial guidance, and uses them to control feature refinement in observed and unobserved regions. This design preserves local visual evidence while exploiting large-scale road and building structure beyond onboard visibility. Experiments on SemanticKITTI and SSCBench-KITTI-360 demonstrate that GeoScene consistently improves both geometric and semantic completion under the geospatial-prior-assisted setting, with the most pronounced benefits for large-scale static and geospatially structured classes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。