arXiv:2506.18798cs.CVcs.AI2025-06被引 3

通过引入物体中心线索,提升视觉3D语义占位预测精度

OC-SOP: Enhancing Vision-Based 3D Semantic Occupancy Prediction by Object-Centric Awareness

  • 用检测分支提取物体级语义线索,融合到占位预测中
  • 在SemanticKITTI上所有类别均达当前最优性能
  • 特别改善动态前景物体的预测效果,适合自动驾驶感知

自动驾驶感知因环境遮挡和场景数据不完整面临重大挑战。为解决此问题,提出语义占位预测(SOP)任务,旨在从图像中联合推断场景的几何结构与语义标签。然而,传统基于摄像头的方法通常对各类别一视同仁,主要依赖局部特征,导致对动态前景物体的预测效果不佳。为此,本文提出物体中心的语义占位预测(OC-SOP),通过检测分支提取高层物体级线索,并将其融入语义占位预测流程。该机制显著提升了前景物体的预测精度,在SemanticKITTI数据集上所有类别均达到当前最优表现。

原文摘要 · Abstract (English)

Autonomous driving perception faces significant challenges due to occlusions and incomplete scene data in the environment. To overcome these issues, the task of semantic occupancy prediction (SOP) is proposed, which aims to jointly infer both the geometry and semantic labels of a scene from images. However, conventional camera-based methods typically treat all categories equally and primarily rely on local features, leading to suboptimal predictions, especially for dynamic foreground objects. To address this, we propose Object-Centric SOP (OC-SOP), a framework that integrates high-level object-centric cues extracted via a detection branch into the semantic occupancy prediction pipeline. This object-centric integration significantly enhances the prediction accuracy for foreground objects and achieves state-of-the-art performance among all categories on SemanticKITTI.

3D语义占位预测自动驾驶物体中心

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。