利用前后帧运动信息提升激光雷达3D目标检测精度
Future Does Matter: Boosting 3D Object Detection with Temporal Motion Estimation in Point Cloud Sequences
- 通过非学习式运动估计生成动态先验,引导跨帧特征融合
- 在Waymo和nuScenes上实现更优的3D检测性能,提升显著
- 适合自动驾驶中需精准时序感知的场景理解任务
准确可靠的激光雷达3D目标检测对自动驾驶中的环境理解至关重要。尽管重要,但点云数据固有的局限性(如远距离与遮挡)仍限制了检测性能。近期研究表明,通过融合多帧视角信息可有效提升检测精度。本文提出新型激光雷达3D检测框架LiSTM,利用跨帧运动预测信息进行时空特征学习。通过非学习型运动估计生成动态先验,设计运动引导特征聚合(MGFA),基于历史与未来帧的物体轨迹,在驾驶序列中构建高斯热图以建模时空相关性,并引导时序特征融合,丰富目标特征表示。同时引入双相关权重模块(DCWM),通过场景与通道级特征抽象,增强过去与未来帧间的交互。最终采用级联交叉注意力解码器优化3D预测。在Waymo与nuScenes数据集上的实验表明,该框架在有效时空特征学习基础上实现了优越的3D检测性能。
原文摘要 · Abstract (English)
Accurate and robust LiDAR 3D object detection is essential for comprehensive scene understanding in autonomous driving. Despite its importance, LiDAR detection performance is limited by inherent constraints of point cloud data, particularly under conditions of extended distances and occlusions. Recently, temporal aggregation has been proven to significantly enhance detection accuracy by fusing multi-frame viewpoint information and enriching the spatial representation of objects. In this work, we introduce a novel LiDAR 3D object detection framework, namely LiSTM, to facilitate spatial-temporal feature learning with cross-frame motion forecasting information. We aim to improve the spatial-temporal interpretation capabilities of the LiDAR detector by incorporating a dynamic prior, generated from a non-learnable motion estimation model. Specifically, Motion-Guided Feature Aggregation (MGFA) is proposed to utilize the object trajectory from previous and future motion states to model spatial-temporal correlations into gaussian heatmap over a driving sequence. This motion-based heatmap then guides the temporal feature fusion, enriching the proposed object features. Moreover, we design a Dual Correlation Weighting Module (DCWM) that effectively facilitates the interaction between past and prospective frames through scene- and channel-wise feature abstraction. In the end, a cascade cross-attention-based decoder is employed to refine the 3D prediction. We have conducted experiments on the Waymo and nuScenes datasets to demonstrate that the proposed framework achieves superior 3D detection performance with effective spatial-temporal feature learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。