首个面向非道路场景的3D语义占位预测基准,填补领域空白
WildOcc: A Benchmark for Off-Road 3D Semantic Occupancy Prediction
- 提出从粗到精的重建流水线生成真实感标注数据
- 构建多模态融合框架,联合图像与点云提升预测精度
- 适合自动驾驶、三维场景理解研究者参考
3D语义占位预测是自动驾驶的关键技术,聚焦于捕捉场景的几何细节。非道路环境蕴含丰富的几何信息,非常适合用于3D语义占位预测任务中的场景重建。然而,现有研究多集中于道路场景,由于缺乏相关数据集和基准,针对非道路环境的3D语义占位预测方法极少。为填补这一空白,本文提出WildOcc,据我们所知是首个提供非道路场景密集占位标注的基准。论文设计了一套真值生成流程,采用粗到精的重建策略以获得更逼真的结果。此外,提出一个融合多帧图像与点云时空信息的多模态3D语义占位预测框架,并引入跨模态知识蒸馏机制,将点云中的几何知识迁移到图像特征中。
原文摘要 · Abstract (English)
3D semantic occupancy prediction is an essential part of autonomous driving, focusing on capturing the geometric details of scenes. Off-road environments are rich in geometric information, therefore it is suitable for 3D semantic occupancy prediction tasks to reconstruct such scenes. However, most of researches concentrate on on-road environments, and few methods are designed for off-road 3D semantic occupancy prediction due to the lack of relevant datasets and benchmarks. In response to this gap, we introduce WildOcc, to our knowledge, the first benchmark to provide dense occupancy annotations for off-road 3D semantic occupancy prediction tasks. A ground truth generation pipeline is proposed in this paper, which employs a coarse-to-fine reconstruction to achieve a more realistic result. Moreover, we introduce a multi-modal 3D semantic occupancy prediction framework, which fuses spatio-temporal information from multi-frame images and point clouds at voxel level. In addition, a cross-modality distillation function is introduced, which transfers geometric knowledge from point clouds to image features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。