用高斯模型融合相机与激光雷达,实现高效精准的3D语义占位预测。
Gaussian Based Adaptive Multi-Modal 3D Semantic Occupancy Prediction
- 基于高斯表示的动态多模态融合,提升感知鲁棒性。
- 计算复杂度线性增长,实测在9.2帧率下稳定运行。
- 适合自动驾驶长尾安全场景,尤其动态环境表现优异。
从稀疏目标检测转向密集3D语义占位预测,是应对自动驾驶长尾安全挑战的必要趋势。然而,现有体素化方法普遍面临计算开销过大、融合过程僵化且在动态环境下易失效的问题。为此,本文提出一种基于高斯的自适应相机-激光雷达多模态3D占位预测模型,通过轻量级3D高斯表示,无缝融合相机的语义优势与激光雷达的几何优势。该方案包含四个关键组件:(1) 激光雷达深度特征聚合(LDFA),采用逐深度可变形采样处理几何稀疏性;(2) 基于熵的特征平滑,利用交叉熵抑制领域特定噪声;(3) 自适应相机-激光雷达融合,根据模型输出动态重校准传感器数据;(4) Gauss-Mamba头,采用选择性状态空间模型实现全局上下文解码,具备线性计算复杂度。实验表明,该模型在nuScenes数据集上达到85.6% mIoU,推理速度达9.2 FPS,显著优于现有方法。
原文摘要 · Abstract (English)
The sparse object detection paradigm shift towards dense 3D semantic occupancy prediction is necessary for dealing with long-tail safety challenges for autonomous vehicles. Nonetheless, the current voxelization methods commonly suffer from excessive computation complexity demands, where the fusion process is brittle, static, and breaks down under dynamic environmental settings. To this end, this research work enhances a novel Gaussian-based adaptive camera-LiDAR multimodal 3D occupancy prediction model that seamlessly bridges the semantic strengths of camera modality with the geometric strengths of LiDAR modality through a memory-efficient 3D Gaussian model. The proposed solution has four key components: (1) LiDAR Depth Feature Aggregation (LDFA), where depth-wise deformable sampling is employed for dealing with geometric sparsity, (2) Entropy-Based Feature Smoothing, where cross-entropy is employed for handling domain-specific noise, (3) Adaptive Camera-LiDAR Fusion, where dynamic recalibration of sensor outputs is performed based on model outputs, and (4) Gauss-Mamba Head that uses Selective State Space Models for global context decoding that enjoys linear computation complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。