arXiv:2508.13601cs.CV2025-08AAAI被引 3

分离语义与几何先验,提升3D场景补全精度

Unleashing Semantic and Geometric Priors for 3D Scene Completion

  • 用双解耦架构分别处理语义和几何信息
  • 在SemanticKITTI上达21.78 mIoU和48.61 IoU
  • 适合自动驾驶与机器人导航的3D感知任务

基于摄像头的3D语义场景补全(SSC)为自动驾驶和机器人导航提供密集的几何与语义感知。然而,现有方法依赖耦合编码器同时传递语义与几何先验,导致模型在相互冲突的需求间权衡,限制整体性能。为此,我们提出FoundationSSC,一种在源级和路径级均实现双解耦的新框架。在源级,引入基础编码器,为语义分支提供丰富语义特征先验,为几何分支提供高保真立体代价体。在路径级,通过专用解耦路径对先验进行优化,生成更优的语义上下文与深度分布。双解耦设计产生解耦且精炼的输入,经混合视图变换生成互补3D特征。此外,提出新型轴向感知融合(AAF)模块,通过各向异性融合解决特征融合难题。大量实验表明,FoundationSSC在语义与几何指标上同步提升,在SemanticKITTI上分别超越前人最优结果+0.23 mIoU和+2.03 IoU;在SSCBench-KITTI-360上达到21.78 mIoU和48.61 IoU的顶尖表现。

原文摘要 · Abstract (English)

Camera-based 3D semantic scene completion (SSC) provides dense geometric and semantic perception for autonomous driving and robotic navigation. However, existing methods rely on a coupled encoder to deliver both semantic and geometric priors, which forces the model to make a trade-off between conflicting demands and limits its overall performance. To tackle these challenges, we propose FoundationSSC, a novel framework that performs dual decoupling at both the source and pathway levels. At the source level, we introduce a foundation encoder that provides rich semantic feature priors for the semantic branch and high-fidelity stereo cost volumes for the geometric branch. At the pathway level, these priors are refined through specialised, decoupled pathways, yielding superior semantic context and depth distributions. Our dual-decoupling design produces disentangled and refined inputs, which are then utilised by a hybrid view transformation to generate complementary 3D features. Additionally, we introduce a novel Axis-Aware Fusion (AAF) module that addresses the often-overlooked challenge of fusing these features by anisotropically merging them into a unified representation. Extensive experiments demonstrate the advantages of FoundationSSC, achieving simultaneous improvements in both semantic and geometric metrics, surpassing prior bests by +0.23 mIoU and +2.03 IoU on SemanticKITTI. Additionally, we achieve state-of-the-art performance on SSCBench-KITTI-360, with 21.78 mIoU and 48.61 IoU.

3D场景补全语义分割几何感知自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。