arXiv:2606.03581cs.CVcs.RO2026-06被引 1

提出新方法提升非结构化场景3D语义占位预测效果

UnsOcc: 3D Semantic Occupancy Prediction in Unstructured Scene via Rendering Fusion

论文配图:UnsOcc: 3D Semantic Occupancy Prediction in Unstructured Scene via Rendering Fusion
图 1 · 摘自论文原文
  • 通过双向渲染监督实现多模态特征对齐
  • 在矿场数据集上mIoU达42.1,优于现有方法
  • 适合自动驾驶在复杂非结构化环境应用

非结构化场景(如露天矿坑)因障碍物不规则、场景稀疏,导致传统3D目标检测失效。3D语义占位预测通过为三维体素分配语义标签,提供密集空间表征,但直接应用于此类场景仍面临挑战:场景稀疏性影响跨模态融合,长尾分布进一步降低预测性能。为此,我们构建了首个面向露天矿坑的专用数据集,并提出UnsOcc框架。核心创新包括:基于渲染的融合模块RenderFusion,通过双向渲染监督增强跨模态特征对齐;以及基于高斯点云的细节感知辅助监督方法GSRefinement,将稀疏3D占位预测投影至稠密2D语义分割图,实现对长尾类别的有效监督。在自建矿场数据集与nuScenes数据集上的大量实验表明,该方法显著优于当前最优方法。

原文摘要 · Abstract (English)

Unstructured scenes present unique challenges for autonomous driving, as irregular obstacles and sparse scene layouts undermine the effectiveness of traditional perception methods such as 3D object detection. 3D semantic occupancy prediction has emerged as a prominent focus due to its ability to provide dense spatial representations by assigning semantic labels to individual voxels in 3D space. However, directly applying 3D semantic occupancy prediction to unstructured scenes remains challenging because scene sparsity hinders effective cross-modal fusion and the more severe long-tail distribution in these scenarios further degrades prediction performance. To validate the effectiveness of our approach, we construct a dedicated dataset of unstructured scenes collected from open-pit mines. Based on this, we propose UnsOcc, a multi-modal 3D semantic occupancy prediction framework that improves robustness in unstructured environments. At its core, we introduce a rendering-based fusion module, RenderFusion, which enhances cross-modal feature alignment through bidirectional rendering supervision. Furthermore, we propose GSRefinement, a detail-aware auxiliary supervision method based on Gaussian Splatting that projects sparse 3D occupancy predictions into dense 2D semantic segmentation maps, enabling effective supervision for long-tail categories. Extensive experiments on both the open-pit mine dataset and the nuScenes dataset demonstrate that our method significantly outperforms existing state-of-the-art approaches.

3D占位自动驾驶多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。