arXiv:2505.10685cs.CV2025-05被引 18

用3D高斯表示融合多模态数据,提升自动驾驶语义占用预测精度与效率

GaussianFormer3D: Multi-Modal Gaussian-based Semantic Occupancy Prediction with 3D Deformable Attention

  • 基于体素初始化3D高斯,利用激光雷达提供几何先验
  • 设计可变形注意力机制,在3D空间融合激光雷达与相机特征
  • 在真实道路与非道路数据集上表现领先,内存消耗更低

3D语义占用预测对实现安全可靠的自动驾驶和机器人导航至关重要。相比仅使用摄像头的感知系统,多模态方法尤其是激光雷达-相机融合方案能生成更精确、更细致的预测结果。尽管体素化场景表示广泛用于语义占用预测,3D高斯作为一种连续且显著更紧凑的替代方案正逐渐兴起。本文提出一种基于多模态3D高斯的语义占用预测框架——GaussianFormer3D,采用3D可变形注意力机制。我们引入体素转高斯初始化策略,利用激光雷达数据为3D高斯提供准确的几何先验,并设计激光雷达引导的3D可变形注意力机制,在升维后的3D空间中融合激光雷达与相机特征以优化高斯表示。在真实道路与非道路自动驾驶数据集上的大量实验表明,GaussianFormer3D在保持更低成本的同时达到当前最优性能。

原文摘要 · Abstract (English)

3D semantic occupancy prediction is essential for achieving safe, reliable autonomous driving and robotic navigation. Compared to camera-only perception systems, multi-modal pipelines, especially LiDAR-camera fusion methods, can produce more accurate and fine-grained predictions. Although voxel-based scene representations are widely used for semantic occupancy prediction, 3D Gaussians have emerged as a continuous and significantly more compact alternative. In this work, we propose a multi-modal Gaussian-based semantic occupancy prediction framework utilizing 3D deformable attention, namely GaussianFormer3D. We introduce a voxel-to-Gaussian initialization strategy that provides 3D Gaussians with accurate geometry priors from LiDAR data, and design a LiDAR-guided 3D deformable attention mechanism to refine these Gaussians using LiDAR-camera fusion features in a lifted 3D space. Extensive experiments on real-world on-road and off-road autonomous driving datasets demonstrate that GaussianFormer3D achieves state-of-the-art prediction performance with reduced memory consumption and improved efficiency.

3D高斯多模态融合语义占用自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。