用卫星与地面图像融合提升3D语义场景补全精度
SGFormer: Satellite-Ground Fusion for 3D Semantic Scene Completion
- 并行编码卫星与地面视角,统一到同一特征空间
- 卫星图像偏差修正使语义补全准确率显著提升
- 自适应权重分配适合多源异构数据融合,适合城市级建模
近期基于相机的场景语义补全(SSC)方法虽在可见区域表现良好,但受限于频繁的视觉遮挡,难以获取完整语义。为此,本文提出首个卫星-地面协同的SSC框架SGFormer,探索卫星与地面图像对在该任务中的潜力。设计双分支结构,平行编码正交的卫星与地面视图,并将其统一至共同特征域。提出地面视图引导策略,在特征编码阶段校正卫星图像偏差,缓解两者间的配准误差。此外,设计自适应加权机制,动态平衡两类视图的贡献。实验表明,SGFormer在SemanticKITTI和SSCBench-KITTI-360数据集上均超越现有最优方法。代码已开源。
原文摘要 · Abstract (English)
Recently, camera-based solutions have been extensively explored for scene semantic completion (SSC). Despite their success in visible areas, existing methods struggle to capture complete scene semantics due to frequent visual occlusions. To address this limitation, this paper presents the first satellite-ground cooperative SSC framework, i.e., SGFormer, exploring the potential of satellite-ground image pairs in the SSC task. Specifically, we propose a dual-branch architecture that encodes orthogonal satellite and ground views in parallel, unifying them into a common domain. Additionally, we design a ground-view guidance strategy that corrects satellite image biases during feature encoding, addressing misalignment between satellite and ground views. Moreover, we develop an adaptive weighting strategy that balances contributions from satellite and ground views. Experiments demonstrate that SGFormer outperforms the state of the art on SemanticKITTI and SSCBench-KITTI-360 datasets. Our code is available on https://github.com/gxytcrc/SGFormer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。