提出高效融合2D与3D特征的框架,提升自动驾驶环境感知精度
OccLoff: Learning Optimized Feature Fusion for 3D Occupancy Prediction
- 设计稀疏融合编码器与熵掩码,直接融合2D图像与LiDAR特征
- 在nuScenes和SemanticKITTI上实现更高精度,计算开销更低
- 适配多种主流模型,提升可迁移性,适合自动驾驶感知研究
3D语义占用预测对精细表征周围环境至关重要,是保障自动驾驶安全的基础。现有基于融合的方法通常需将图像特征进行2D到3D视图变换,再通过高成本的3D操作融合LiDAR特征,导致计算开销大、精度受限。同时,当前研究多聚焦特定网络架构设计,忽视更基础的语义特征学习,限制了方法的可迁移性。为此,我们提出OccLoff框架,通过稀疏融合编码器与熵掩码,直接融合3D与2D特征,在提升准确率的同时降低计算负担。此外,引入可迁移的代理损失函数与自适应困难样本加权算法,显著增强多个前沿方法的性能。在nuScenes与SemanticKITTI基准上的大量实验验证了框架优势,消融实验证明各模块有效性。
原文摘要 · Abstract (English)
3D semantic occupancy prediction is crucial for finely representing the surrounding environment, which is essential for ensuring the safety in autonomous driving. Existing fusion-based occupancy methods typically involve performing a 2D-to-3D view transformation on image features, followed by computationally intensive 3D operations to fuse these with LiDAR features, leading to high computational costs and reduced accuracy. Moreover, current research on occupancy prediction predominantly focuses on designing specific network architectures, often tailored to particular models, with limited attention given to the more fundamental aspect of semantic feature learning. This gap hinders the development of more transferable methods that could enhance the performance of various occupancy models. To address these challenges, we propose OccLoff, a framework that Learns to Optimize Feature Fusion for 3D occupancy prediction. Specifically, we introduce a sparse fusion encoder with entropy masks that directly fuses 3D and 2D features, improving model accuracy while reducing computational overhead. Additionally, we propose a transferable proxy-based loss function and an adaptive hard sample weighting algorithm, which enhance the performance of several state-of-the-art methods. Extensive evaluations on the nuScenes and SemanticKITTI benchmarks demonstrate the superiority of our framework, and ablation studies confirm the effectiveness of each proposed module.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。