提出SEGT框架,提升激光雷达目标检测的精度与效率。
SEGT: A General Spatial Expansion Group Transformer for nuScenes Lidar-based Object Detection Task
- 用空间扩展分组注意力机制处理点云稀疏性问题
- 在nuScenes上达73.9(无TTA)和74.5(有TTA)NDS分数
- 适合关注自动驾驶感知与点云建模的研究者
本文提出一种基于Transformer的新型框架——空间扩展分组变换器(SEGT),用于nuScenes激光雷达目标检测任务。为高效处理点云的不规则与稀疏特性,该方法将体素迁移到具有通用空间扩展策略的独立有序场中,并采用分组注意力机制提取各场内的专属特征图。随后,通过交替应用多种扩展策略,融合不同有序场间的特征表示,增强模型对全局空间信息的捕捉能力。在nuScenes激光雷达目标检测测试集上评估,未使用测试时增强(TTA)时达到73.9的NDS分数,使用TTA时达74.5,显著优于现有方法。本方法在nuScenes激光雷达目标检测任务中排名第一。
原文摘要 · Abstract (English)
In the technical report, we present a novel transformer-based framework for nuScenes lidar-based object detection task, termed Spatial Expansion Group Transformer (SEGT). To efficiently handle the irregular and sparse nature of point cloud, we propose migrating the voxels into distinct specialized ordered fields with the general spatial expansion strategies, and employ group attention mechanisms to extract the exclusive feature maps within each field. Subsequently, we integrate the feature representations across different ordered fields by alternately applying diverse expansion strategies, thereby enhancing the model's ability to capture comprehensive spatial information. The method was evaluated on the nuScenes lidar-based object detection test dataset, achieving an NDS score of 73.9 without Test-Time Augmentation (TTA) and 74.5 with TTA, demonstrating the effectiveness of the proposed method. Notably, our method ranks the 1st place in the nuScenes lidar-based object detection task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。