用3D高斯表示融合多传感器数据,提升自动驾驶环境占位预测精度与效率。
GaussianFusionOcc: A Seamless Sensor Fusion Approach for 3D Occupancy Prediction Using 3D Gaussians
- 采用3D高斯表示替代密集网格,实现更高效的环境建模。
- 通过可变形注意力机制融合相机、激光雷达和雷达数据,提升占位预测精度。
- 在多种传感器组合下表现优异,适合追求实时性与准确性的自动驾驶系统。
3D语义占位预测是自动驾驶中的关键任务,能实现对复杂环境的精确与安全理解与导航。可靠预测依赖于有效的多传感器融合,因不同模态包含互补信息。与依赖密集网格表示的传统方法不同,本文提出的GaussianFusionOcc采用语义3D高斯表示,并引入创新的传感器融合机制。通过无缝融合相机、激光雷达和雷达数据,实现更精准且可扩展的占位预测,同时3D高斯表示显著提升内存效率与推理速度。GaussianFusionOcc使用模态无关的可变形注意力机制提取各传感器的关键特征,并用于优化高斯属性,从而获得更准确的环境表征。在多种传感器组合下的大量测试验证了该方法的通用性。通过结合多模态融合的鲁棒性与高斯表示的高效性,GaussianFusionOcc超越当前最先进模型。
原文摘要 · Abstract (English)
3D semantic occupancy prediction is one of the crucial tasks of autonomous driving. It enables precise and safe interpretation and navigation in complex environments. Reliable predictions rely on effective sensor fusion, as different modalities can contain complementary information. Unlike conventional methods that depend on dense grid representations, our approach, GaussianFusionOcc, uses semantic 3D Gaussians alongside an innovative sensor fusion mechanism. Seamless integration of data from camera, LiDAR, and radar sensors enables more precise and scalable occupancy prediction, while 3D Gaussian representation significantly improves memory efficiency and inference speed. GaussianFusionOcc employs modality-agnostic deformable attention to extract essential features from each sensor type, which are then used to refine Gaussian properties, resulting in a more accurate representation of the environment. Extensive testing with various sensor combinations demonstrates the versatility of our approach. By leveraging the robustness of multi-modal fusion and the efficiency of Gaussian representation, GaussianFusionOcc outperforms current state-of-the-art models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。