arXiv:2601.22729cs.CV2026-01被引 6

用高斯表示融合相机与激光雷达,提升3D语义占据预测精度与鲁棒性

GaussianOcc3D: A Gaussian-Based Adaptive Multi-modal 3D Occupancy Prediction

  • 以连续高斯点云表示场景,高效融合相机语义与激光雷达几何信息
  • 在Occ3D、SurroundOcc、SemanticKITTI上分别达到49.4%、28.9%、25.2%的mIoU
  • 对雨夜等复杂场景表现稳健,适合自动驾驶环境理解任务

3D语义占据预测是自动驾驶中的关键任务,提供对周围环境密集且细粒度的理解。然而,单模态方法在相机语义与激光雷达几何之间存在权衡。现有多模态框架常面临模态异质性、空间错位及表示困境——体素计算开销大,鸟瞰图(BEV)表示有损。我们提出GaussianOcc3D,一种基于高斯的多模态框架,通过内存高效的连续3D高斯表示连接相机与激光雷达。引入四个模块:(1) 激光雷达深度特征聚合(LDFA),使用深度可变形采样将稀疏信号映射到高斯原语;(2) 基于熵的特征平滑(EBFS),缓解领域噪声;(3) 自适应相机-激光雷达融合(ACLF),通过不确定性感知重加权提升传感器可靠性;(4) Gauss-Mamba Head,利用选择性状态空间模型实现线性复杂度的全局上下文建模。在Occ3D、SurroundOcc和SemanticKITTI基准上评估,取得49.4%、28.9%、25.2%的mIoU,表现出对雨夜等挑战性条件的优异鲁棒性。

原文摘要 · Abstract (English)

3D semantic occupancy prediction is a pivotal task in autonomous driving, providing a dense and fine-grained understanding of the surrounding environment, yet single-modality methods face trade-offs between camera semantics and LiDAR geometry. Existing multi-modal frameworks often struggle with modality heterogeneity, spatial misalignment, and the representation crisis--where voxels are computationally heavy and BEV alternatives are lossy. We present GaussianOcc3D, a multi-modal framework bridging camera and LiDAR through a memory-efficient, continuous 3D Gaussian representation. We introduce four modules: (1) LiDAR Depth Feature Aggregation (LDFA), using depth-wise deformable sampling to lift sparse signals onto Gaussian primitives; (2) Entropy-Based Feature Smoothing (EBFS) to mitigate domain noise; (3) Adaptive Camera-LiDAR Fusion (ACLF) with uncertainty-aware reweighting for sensor reliability; and (4) a Gauss-Mamba Head leveraging Selective State Space Models for global context with linear complexity. Evaluations on Occ3D, SurroundOcc, and SemanticKITTI benchmarks demonstrate state-of-the-art performance, achieving mIoU scores of 49.4%, 28.9%, and 25.2% respectively. GaussianOcc3D exhibits superior robustness across challenging rainy and nighttime conditions.

3D占据预测多模态融合高斯表示自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。