用物体为中心的占据表示,提升3D检测精度与细节还原能力
Towards Flexible 3D Perception: Object-Centric Occupancy Completion Augments 3D Object Detection
- 以物体为中心构建高分辨率占据图,仅关注前景物体区域
- 在Waymo数据集上显著提升远距离和不完整物体的检测效果
- 适用于自动驾驶中复杂场景的精准感知,尤其适合改进现有检测器
尽管3D目标边界框广泛用于自动驾驶感知,但难以捕捉物体的精确几何细节。近年来,占据表示成为3D场景感知的有力替代方案。然而,受限于计算资源,大场景下构建高分辨率占据图仍不可行。鉴于前景物体仅占场景一小部分,本文提出物体为中心的占据表示作为边界框的补充,不仅提供更精细的物体细节,还可在实际应用中实现更高体素分辨率。我们在数据与算法两方面推进该方向:一方面,构建首个从零开始的物体为中心占据数据集;另一方面,提出一种新型物体为中心占据补全网络,配备隐式形状解码器,可处理动态尺寸占据生成,并利用长序列时间信息,准确补全不准确的物体提议。实验表明,该方法在噪声检测与跟踪条件下仍具鲁棒性。此外,其占据特征显著提升主流3D检测器性能,尤其在Waymo Open Dataset中对不完整或远距离物体效果突出。
原文摘要 · Abstract (English)
While 3D object bounding box (bbox) representation has been widely used in autonomous driving perception, it lacks the ability to capture the precise details of an object's intrinsic geometry. Recently, occupancy has emerged as a promising alternative for 3D scene perception. However, constructing a high-resolution occupancy map remains infeasible for large scenes due to computational constraints. Recognizing that foreground objects only occupy a small portion of the scene, we introduce object-centric occupancy as a supplement to object bboxes. This representation not only provides intricate details for detected objects but also enables higher voxel resolution in practical applications. We advance the development of object-centric occupancy perception from both data and algorithm perspectives. On the data side, we construct the first object-centric occupancy dataset from scratch using an automated pipeline. From the algorithmic standpoint, we introduce a novel object-centric occupancy completion network equipped with an implicit shape decoder that manages dynamic-size occupancy generation. This network accurately predicts the complete object-centric occupancy volume for inaccurate object proposals by leveraging temporal information from long sequences. Our method demonstrates robust performance in completing object shapes under noisy detection and tracking conditions. Additionally, we show that our occupancy features significantly enhance the detection results of state-of-the-art 3D object detectors, especially for incomplete or distant objects in the Waymo Open Dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。