提出垂直切片融合框架,提升3D语义占位的高精度感知能力
SliceSemOcc: Vertical Slice Based Multimodal 3D Semantic Occupancy Representation
- 按高度轴分层提取特征,结合全局与局部切片信息
- 在nuScenes数据集上平均交并比显著提升,小物体识别效果更优
- 适合自动驾驶中需精细垂直结构理解的场景
为满足自动驾驶对高精度三维感知的需求,三维语义占位预测已成为关键研究方向。与局限于二维平面的鸟瞰图方法不同,占位预测利用完整的三维体素网格,在所有维度上建模空间结构,从而捕捉垂直方向上的语义变化。然而,现有方法大多在处理体素特征时忽略高度轴信息,传统SENet式通道注意力对所有高度层赋予相同权重,限制了对不同高度特征的关注能力。为此,我们提出SliceSemOcc,一种基于垂直切片的多模态三维语义占位表示框架。具体而言,通过全局和局部垂直切片提取体素特征,并设计全局-局部融合模块,自适应融合细粒度空间细节与整体上下文信息。此外,提出SEAttention3D模块,通过平均池化保持高度分辨率,并为每层高度动态分配通道注意力权重。在nuScenes-SurroundOcc和nuScenes-OpenOccupancy数据集上的大量实验表明,该方法显著提升了平均交并比,尤其在多数小物体类别上表现突出。详细的消融实验证实了SliceSemOcc框架的有效性。
原文摘要 · Abstract (English)
Driven by autonomous driving's demands for precise 3D perception, 3D semantic occupancy prediction has become a pivotal research topic. Unlike bird's-eye-view (BEV) methods, which restrict scene representation to a 2D plane, occupancy prediction leverages a complete 3D voxel grid to model spatial structures in all dimensions, thereby capturing semantic variations along the vertical axis. However, most existing approaches overlook height-axis information when processing voxel features. And conventional SENet-style channel attention assigns uniform weight across all height layers, limiting their ability to emphasize features at different heights. To address these limitations, we propose SliceSemOcc, a novel vertical slice based multimodal framework for 3D semantic occupancy representation. Specifically, we extract voxel features along the height-axis using both global and local vertical slices. Then, a global local fusion module adaptively reconciles fine-grained spatial details with holistic contextual information. Furthermore, we propose the SEAttention3D module, which preserves height-wise resolution through average pooling and assigns dynamic channel attention weights to each height layer. Extensive experiments on nuScenes-SurroundOcc and nuScenes-OpenOccupancy datasets verify that our method significantly enhances mean IoU, achieving especially pronounced gains on most small-object categories. Detailed ablation studies further validate the effectiveness of the proposed SliceSemOcc framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。