arXiv:2511.13031cs.CV2025-11AAAI

提出面向物体的3D语义场景补全方法,提升复杂环境下的预测精度。

Towards 3D Object-Centric Feature Learning for Semantic Scene Completion

  • 将场景分解为独立物体实例,实现以物体为中心的特征学习。
  • 在SemanticKITTI和SSCBench-KITTI360上分别达到17.40和20.28的mIoU。
  • 适用于自动驾驶中需要精细语义理解的复杂场景补全任务。

基于视觉的3D语义场景补全(SSC)因其在自动驾驶中的潜力而受到越来越多关注。现有方法多采用以自我为中心的范式,在整个场景中聚合与传播特征,但常忽略细粒度的物体级细节,导致复杂环境中语义与几何模糊。为此,我们提出Ocean,一种以物体为中心的预测框架,将场景分解为独立物体实例,以实现更精确的语义占据预测。首先,使用轻量级分割模型MobileSAM从输入图像中提取实例掩码;其次,引入3D语义分组注意力模块,利用线性注意力在3D空间中聚合物体中心特征;为处理分割错误与缺失实例,设计全局相似性引导注意力模块,利用分割特征进行全局交互;最后,提出实例感知局部扩散模块,通过生成过程改进实例特征,并在鸟瞰图(BEV)空间中优化场景表示。在SemanticKITTI和SSCBench-KITTI360基准上的大量实验表明,Ocean取得当前最优性能,mIoU分别为17.40和20.28。

原文摘要 · Abstract (English)

Vision-based 3D Semantic Scene Completion (SSC) has received growing attention due to its potential in autonomous driving. While most existing approaches follow an ego-centric paradigm by aggregating and diffusing features over the entire scene, they often overlook fine-grained object-level details, leading to semantic and geometric ambiguities, especially in complex environments. To address this limitation, we propose Ocean, an object-centric prediction framework that decomposes the scene into individual object instances to enable more accurate semantic occupancy prediction. Specifically, we first employ a lightweight segmentation model, MobileSAM, to extract instance masks from the input image. Then, we introduce a 3D Semantic Group Attention module that leverages linear attention to aggregate object-centric features in 3D space. To handle segmentation errors and missing instances, we further design a Global Similarity-Guided Attention module that leverages segmentation features for global interaction. Finally, we propose an Instance-aware Local Diffusion module that improves instance features through a generative process and subsequently refines the scene representation in the BEV space. Extensive experiments on the SemanticKITTI and SSCBench-KITTI360 benchmarks demonstrate that Ocean achieves state-of-the-art performance, with mIoU scores of 17.40 and 20.28, respectively.

3D语义补全物体中心自动驾驶特征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。