arXiv:2411.13628cs.CV2024-11被引 4

用状态空间模型高效融合时序信息,提升多视角3D目标检测精度

MambaDETR: Query-based Temporal Modeling using State Space Model for Multi-View 3D Object Detection

  • 基于状态空间模型实现线性复杂度的时序融合,避免传统Transformer的计算瓶颈
  • 在nuScenes数据集上达到当前最优性能,显著优于现有时序融合方法
  • 设计运动消除模块,过滤静态物体干扰,增强动态目标检测能力

利用时序信息提升自动驾驶中3D目标检测性能近年来取得显著进展。传统基于Transformer的时序融合方法在序列长度增加时面临二次计算开销和信息衰减问题。本文提出一种新方法MambaDETR,通过高效的状态空间实现时序融合,并设计运动消除模块以移除相对静止物体,优化时序信息利用。在标准nuScenes基准上,所提MambaDETR在3D目标检测任务中表现优异,成为现有时序融合方法中的最先进水平。

原文摘要 · Abstract (English)

Utilizing temporal information to improve the performance of 3D detection has made great progress recently in the field of autonomous driving. Traditional transformer-based temporal fusion methods suffer from quadratic computational cost and information decay as the length of the frame sequence increases. In this paper, we propose a novel method called MambaDETR, whose main idea is to implement temporal fusion in the efficient state space. Moreover, we design a Motion Elimination module to remove the relatively static objects for temporal fusion. On the standard nuScenes benchmark, our proposed MambaDETR achieves remarkable result in the 3D object detection task, exhibiting state-of-the-art performance among existing temporal fusion methods.

3D检测时序建模状态空间自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。