针对3D语义占位中的多模态融合难题,提出自适应跨模态对齐与体渲染优化方法。
TACOcc:Target-Adaptive Cross-Modal Fusion with Volume Rendering for 3D Semantic Occupancy
- 根据目标尺度动态调整融合邻域,提升跨模态特征对齐精度。
- 在nuScenes和SemanticKITTI上,3D语义占位准确率显著优于现有方法。
- 适合自动驾驶场景下高精度3D环境感知的研究与应用。
多模态3D占位预测性能受限于无效融合,主要源于固定融合策略引发的几何-语义不匹配,以及稀疏噪声标注导致的表面细节丢失。不匹配源于点云与图像特征在尺度和分布上的异质性,导致固定邻域融合存在偏差。为此,我们提出一种目标尺度自适应的双向对称检索机制,对大目标扩大邻域以增强上下文感知,对小目标缩小邻域以提高效率并抑制噪声,实现精准的跨模态特征对齐。该机制显式建立空间对应关系,提升融合准确性。针对表面细节丢失问题,稀疏标签监督不足,影响小物体预测。我们引入基于3D Gaussian Splatting的改进体渲染流程,以融合特征为输入进行图像渲染,施加光度一致性监督,并联合优化2D-3D一致性,增强表面细节重建同时抑制噪声传播。综上,我们提出TACOcc框架,通过自适应多模态融合与体渲染监督,显著提升3D语义占位性能。在nuScenes和SemanticKITTI基准上实验验证了其有效性。
原文摘要 · Abstract (English)
The performance of multi-modal 3D occupancy prediction is limited by ineffective fusion, mainly due to geometry-semantics mismatch from fixed fusion strategies and surface detail loss caused by sparse, noisy annotations. The mismatch stems from the heterogeneous scale and distribution of point cloud and image features, leading to biased matching under fixed neighborhood fusion. To address this, we propose a target-scale adaptive, bidirectional symmetric retrieval mechanism. It expands the neighborhood for large targets to enhance context awareness and shrinks it for small ones to improve efficiency and suppress noise, enabling accurate cross-modal feature alignment. This mechanism explicitly establishes spatial correspondences and improves fusion accuracy. For surface detail loss, sparse labels provide limited supervision, resulting in poor predictions for small objects. We introduce an improved volume rendering pipeline based on 3D Gaussian Splatting, which takes fused features as input to render images, applies photometric consistency supervision, and jointly optimizes 2D-3D consistency. This enhances surface detail reconstruction while suppressing noise propagation. In summary, we propose TACOcc, an adaptive multi-modal fusion framework for 3D semantic occupancy prediction, enhanced by volume rendering supervision. Experiments on the nuScenes and SemanticKITTI benchmarks validate its effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。