arXiv:2503.10579cs.CV2025-03ICRA被引 1

通过语义监督融合时空信息,提升点云3D目标检测精度

Semantic-Supervised Spatial-Temporal Fusion for LiDAR-based 3D Object Detection

  • 设计时空融合模块,缓解物体运动导致的点云空间错位
  • 在nuScenes上实现NDS提升约2.8%,通用性强
  • 利用点级语义标签增强稀疏数据,适合自动驾驶感知任务

基于激光雷达的3D目标检测因点云固有的稀疏性面临挑战。常用方法是利用长时间序列的激光雷达数据来丰富输入。然而,高效利用时空信息仍是开放问题。本文提出一种新型语义监督时空融合(ST-Fusion)方法,引入新融合模块以缓解物体运动引起的时空错位,并通过特征级语义监督充分挖掘该模块潜力。具体而言,ST-Fusion包含空间聚合(SA)模块和时间合并(TM)模块。SA模块采用感受野逐层扩大的卷积层,从局部区域聚合物体特征以缓解空间错位;TM模块基于注意力机制动态提取前序帧中的物体特征,实现全面的时序表征。此外,在语义监督中,提出语义注入方法,通过注入点级语义标签丰富稀疏激光雷达数据,用于训练教师模型,并以提出的物体感知损失提供特征级重建目标。在多个激光雷达检测器上的大量实验表明,本方法有效且具有普适性,在nuScenes基准上实现了约+2.8%的NDS提升。

原文摘要 · Abstract (English)

LiDAR-based 3D object detection presents significant challenges due to the inherent sparsity of LiDAR points. A common solution involves long-term temporal LiDAR data to densify the inputs. However, efficiently leveraging spatial-temporal information remains an open problem. In this paper, we propose a novel Semantic-Supervised Spatial-Temporal Fusion (ST-Fusion) method, which introduces a novel fusion module to relieve the spatial misalignment caused by the object motion over time and a feature-level semantic supervision to sufficiently unlock the capacity of the proposed fusion module. Specifically, the ST-Fusion consists of a Spatial Aggregation (SA) module and a Temporal Merging (TM) module. The SA module employs a convolutional layer with progressively expanding receptive fields to aggregate the object features from the local regions to alleviate the spatial misalignment, the TM module dynamically extracts object features from the preceding frames based on the attention mechanism for a comprehensive sequential presentation. Besides, in the semantic supervision, we propose a Semantic Injection method to enrich the sparse LiDAR data via injecting the point-wise semantic labels, using it for training a teacher model and providing a reconstruction target at the feature level supervised by the proposed object-aware loss. Extensive experiments on various LiDAR-based detectors demonstrate the effectiveness and universality of our proposal, yielding an improvement of approximately +2.8% in NDS based on the nuScenes benchmark.

3D检测点云处理时空融合自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。