arXiv:2510.13198cs.CV2025-10ICRA被引 1

通过多层级特征融合提升摄像头驱动的占位预测精度

Complementary Information Guided Occupancy Prediction via Multi-Level Representation Fusion

论文配图:Complementary Information Guided Occupancy Prediction via Multi-Level Representation Fusion
图 1 · 摘自论文原文
  • 融合分割、图形与深度三类多层级特征,增强视觉表征
  • 在SemanticKITTI上达到当前最优性能,无需增加训练成本
  • 适合自动驾驶3D感知研究者及系统开发者参考

基于摄像头的占位预测是自动驾驶中主流的三维感知方法,旨在从二维图像中推断完整的三维场景几何与语义信息。现有方法多通过结构改进(如轻量级主干网络和复杂级联框架)提升性能,虽有成效但受限。少有研究关注表示融合,导致2D图像中丰富的特征多样性未被充分利用。为此,我们提出一种两阶段占位预测框架CIGOcc,通过提取输入图像的分割、图形与深度特征,并引入可变形多层级融合机制进行特征融合;同时,利用SAM模型的知识蒸馏进一步提升预测精度。该方法在不增加训练成本的前提下,在SemanticKITTI基准上实现当前最优性能。代码见补充材料,后续将发布于https://github.com/VitaLemonTea1/CIGOcc。

原文摘要 · Abstract (English)

Camera-based occupancy prediction is a mainstream approach for 3D perception in autonomous driving, aiming to infer complete 3D scene geometry and semantics from 2D images. Almost existing methods focus on improving performance through structural modifications, such as lightweight backbones and complex cascaded frameworks, with good yet limited performance. Few studies explore from the perspective of representation fusion, leaving the rich diversity of features in 2D images underutilized. Motivated by this, we propose \textbf{CIGOcc, a two-stage occupancy prediction framework based on multi-level representation fusion. \textbf{CIGOcc extracts segmentation, graphics, and depth features from an input image and introduces a deformable multi-level fusion mechanism to fuse these three multi-level features. Additionally, CIGOcc incorporates knowledge distilled from SAM to further enhance prediction accuracy. Without increasing training costs, CIGOcc achieves state-of-the-art performance on the SemanticKITTI benchmark. The code is provided in the supplementary material and will be released https://github.com/VitaLemonTea1/CIGOcc

占位预测多模态融合自动驾驶图像理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。