arXiv:2606.31688cs.CV2026-06

单帧点云实现高精度语义占位预测,效率提升2.1倍

Semantic Occupancy Prediction with Dual Range-Voxel Representation

论文配图:Semantic Occupancy Prediction with Dual Range-Voxel Representation
图 1 · 摘自论文原文
  • 用双视角表征融合距离图与体素几何信息
  • nuScenes上mIoU提升5.4%,推理速度加快2.1倍
  • 适合追求高效实时的自动驾驶感知系统

基于激光雷达的3D语义占位预测对自动驾驶系统至关重要。由于点云存在稀疏性和不完整性,现有方法常采用多帧堆叠以获取稠密空间信息,但带来计算开销大和位姿噪声等问题。本文提出双范围-体素表示(DRVR),仅用单帧点云,通过距离图编码器提取场景紧凑上下文,设计几何感知体素编码器分层提取多尺度体素特征并融合,再通过体素-距离双向融合模块协同两类特征。在nuScenes-Occupancy、SemanticKITTI和SemanticPOSS上的实验表明,该方法显著优于现有方法。尤其在nuScenes-Occupancy上,单帧DRVR相比多帧方法实现5.4%的mIoU提升,并达到2.1倍加速。

原文摘要 · Abstract (English)

LiDAR-based 3D semantic occupancy prediction, which aims to provide accurate and comprehensive scene representation, is crucial for autonomous driving systems. As point clouds suffer from sparsity and incompleteness, leading to insufficient semantic learning and difficult occupancy perception, existing methods often stack multi-sweep point clouds to obtain dense spatial information. However, such a naive strategy also results in efficiency (e.g., additional computational burden) and robustness (e.g., pose transformation noise) concerns, which hinder their practical applications. In this work, we propose a Dual Range-Voxel Representation (DRVR) that leverages the range-view context and voxel-view geometry of single-sweep point clouds for 3D semantic occupancy prediction, eliminating the concerns associated with the multi-sweeps. Specifically, we use the range-view encoder to extract the compact context of the scene. To fully exploit the spatial information, we design a geometry-aware voxel-view encoder that extracts multi-scale voxel-view features separately and combines them for better geometric occupancy prediction. Moreover, we propose a range-voxel fusion module to cooperate range- and voxel-view features via voxel-to-range and range-to-voxel fusions. Extensive experiments on nuScenes-Occupancy, SemanticKITTI and SemanticPOSS show the superiority of our method. Especially on nuScenes-Occupancy, our single-sweep DRVR achieves 5.4% improvement in mIoU and 2.1x acceleration compared to the multi-sweep method.

语义占位激光雷达单帧预测自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。