用稀疏高斯表示3D占用,提升多模态感知效率与精度
Gau-Occ: Geometry-Completed Gaussians for Multi-Modal 3D Occupancy Prediction
- 用3D高斯替代密集体素,大幅降低计算开销
- 通过激光雷达补全扩散器恢复缺失结构,初始化稳定高斯锚点
- 多视角图像通过几何对齐采样融合,适合自动驾驶场景
3D语义占用预测对自动驾驶至关重要。尽管多模态融合相比纯视觉方法提升了精度,但通常依赖计算成本高昂的密集体素或鸟瞰图张量。我们提出 Gau-Occ,一种无需密集体积处理的多模态框架,将场景建模为紧凑的语义3D高斯集合。为确保几何完整性,提出激光雷达补全扩散器(LCD),从稀疏激光雷达中恢复缺失结构,以初始化鲁棒的高斯锚点。此外,引入高斯锚点融合(GAF),通过几何对齐的2D采样和跨模态对齐,高效整合多视角图像语义。通过优化这些紧凑的高斯描述符,Gau-Occ 同时捕捉空间一致性与语义区分性。在多个挑战性基准上的实验表明,Gau-Occ 在保持显著计算效率的同时达到当前最优性能。
原文摘要 · Abstract (English)
3D semantic occupancy prediction is crucial for autonomous driving. While multi-modal fusion improves accuracy over vision-only methods, it typically relies on computationally expensive dense voxel or BEV tensors. We present Gau-Occ, a multi-modal framework that bypasses dense volumetric processing by modeling the scene as a compact collection of semantic 3D Gaussians. To ensure geometric completeness, we propose a LiDAR Completion Diffuser (LCD) that recovers missing structures from sparse LiDAR to initialize robust Gaussian anchors. Furthermore, we introduce Gaussian Anchor Fusion (GAF), which efficiently integrates multi-view image semantics via geometry-aligned 2D sampling and cross-modal alignment. By refining these compact Gaussian descriptors, Gau-Occ captures both spatial consistency and semantic discriminability. Extensive experiments across challenging benchmarks demonstrate that Gau-Occ achieves state-of-the-art performance with significant computational efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。