用3D高斯表示统一多模态感知,提升细节保留与融合效率
GaussianFusion: Unified 3D Gaussian Representation for Multi-Modal Fusion Perception

- 以连续3D高斯空间替代离散网格,实现多模态特征统一
- 在nuScenes上检测性能超BEVFusion 2.6 NDS,语义占据提升1.55 mIoU
- 通用框架支持多种任务,仅用30%高斯数却快4.5倍
鸟瞰图(BEV)表示法将多传感器特征融合于统一空间,是实现全面3D感知的主要方法。然而,BEV的离散网格表示会造成显著细节丢失,并限制特征对齐与跨模态信息交互。本文突破传统BEV范式,提出基于3D高斯表示的多模态融合新框架。该方法在共享连续3D高斯空间中自然统一多模态特征,有效保留边缘和精细纹理细节。为此,设计了基于前向投影的多模态高斯初始化模块和共享跨模态高斯编码器,通过注意力机制迭代更新高斯属性。GaussianFusion为天然任务无关模型,其统一高斯表示可自然支持多种3D感知任务。大量实验表明其通用性与鲁棒性:在nuScenes数据集上,其3D物体检测性能超越基线BEVFusion 2.6 NDS;其变体在3D语义占用任务上相比GaussFormer提升1.55 mIoU,仅使用30%高斯数量,速度提升450%。
原文摘要 · Abstract (English)
The bird's-eye view (BEV) representation enables multi-sensor features to be fused within a unified space, serving as the primary approach for achieving comprehensive 3D perception. However, the discrete grid representation of BEV leads to significant detail loss and limits feature alignment and cross-modal information interaction in multimodal fusion perception. In this work, we break from the conventional BEV paradigm and propose a new universal framework for multi-modal fusion based on 3D Gaussian representation. This approach naturally unifies multi-modal features within a shared and continuous 3D Gaussian space, effectively preserving edge and fine texture details. To achieve this, we design a novel forward-projection-based multi-modal Gaussian initialization module and a shared cross-modal Gaussian encoder that iteratively updates Gaussian properties based on an attention mechanism. GaussianFusion is inherently a task-agnostic model, with its unified Gaussian representation naturally supporting various 3D perception tasks. Extensive experiments demonstrate the generality and robustness of GaussianFusion. On the nuScenes dataset, it outperforms the 3D object detection baseline BEVFusion by 2.6 NDS. Its variant surpasses GaussFormer on 3D semantic occupancy with 1.55 mIoU improvement while using only 30% of the Gaussians and achieving a 450% speedup.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。