arXiv:2410.07758cs.CV2024-10被引 2

提升路边视角3D检测精度,通过高度对齐优化特征投影

HeightFormer: A Semantic Alignment Monocular 3D Object Detection Method from Roadside Perspective

  • 引入空间注意力与体素池化架构,增强2D到3D特征映射
  • 在Rope3D和DAIR-V2X-I上车辆与骑行者检测性能领先
  • 适合智能交通系统中车路协同场景的高精度感知需求

车载3D目标检测技术作为自动驾驶关键技术受到广泛关注,但针对路边传感器在3D交通目标检测中的应用研究较少。现有方法基于视锥体进行高度估计实现2D图像特征到3D特征的投影,但未考虑高度对齐及鸟瞰图特征提取效率。本文提出一种新型3D目标检测框架,融合Spatial Former与Voxel Pooling Former,改进基于高度估计的2D-to-3D投影。在Rope3D和DAIR-V2X-I数据集上开展大量实验,结果表明该算法在车辆与骑行者检测任务中均表现更优,具备良好鲁棒性与泛化能力。提升路边3D检测精度有助于构建安全可信的车路协同智能交通系统,推动自动驾驶大规模应用。代码与预训练模型将公开于https://anonymous.4open.science/r/HeightFormer。

原文摘要 · Abstract (English)

The on-board 3D object detection technology has received extensive attention as a critical technology for autonomous driving, while few studies have focused on applying roadside sensors in 3D traffic object detection. Existing studies achieve the projection of 2D image features to 3D features through height estimation based on the frustum. However, they did not consider the height alignment and the extraction efficiency of bird's-eye-view features. We propose a novel 3D object detection framework integrating Spatial Former and Voxel Pooling Former to enhance 2D-to-3D projection based on height estimation. Extensive experiments were conducted using the Rope3D and DAIR-V2X-I dataset, and the results demonstrated the outperformance of the proposed algorithm in the detection of both vehicles and cyclists. These results indicate that the algorithm is robust and generalized under various detection scenarios. Improving the accuracy of 3D object detection on the roadside is conducive to building a safe and trustworthy intelligent transportation system of vehicle-road coordination and promoting the large-scale application of autonomous driving. The code and pre-trained models will be released on https://anonymous.4open.science/r/HeightFormer.

3D检测车路协同视觉感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。