用Mamba2提升车载感知中大物体的3D检测精度
MambaBEV: An EV-based 3D detection model with Mamba2
- 用Mamba2构建时序融合模块,增强鸟瞰图全局上下文建模
- 在nuScenes上达到51.7% NDS和42.7% mAP,优于传统方法
- 适合需要长序列建模的自动驾驶端到端系统
自动驾驶中的精准3D目标检测依赖于鸟瞰图(BEV)感知与有效的时序融合。然而,现有基于卷积或可变形自注意力的融合策略难以在BEV空间建模全局上下文,导致大物体检测精度下降。为此,我们提出MambaBEV,一种基于Mamba2(一种优化的长序列处理状态空间模型)的新型BEV 3D检测模型。核心贡献是TemporalMamba时序融合模块,通过针对序列处理设计的BEV特征离散重排机制增强全局上下文建模;同时引入基于Mamba的DETR头以改进多目标表征。在nuScenes数据集上的评估显示,MambaBEV-base达到51.7% NDS和42.7% mAP。此外,在端到端自动驾驶范式下的评估验证了其在运动预测与规划中的有效性。结果表明,状态空间模型在提升自动驾驶感知系统中全局上下文理解与大物体检测方面具有潜力。
原文摘要 · Abstract (English)
Accurate 3D object detection in autonomous driving relies on Bird's Eye View (BEV) perception and effective temporal fusion. However, existing fusion strategies based on convolutional layers or deformable self-attention struggle to model global context in BEV space, leading to reduced accuracy for large objects.To address this limitation, we propose MambaBEV, a novel BEV-based 3D object detection model that leverages Mamba2, an advanced state-space model (SSM) optimized for long-sequence processing. Our key contribution is TemporalMamba, a temporal fusion module that enhances global context modeling through a BEV feature discrete rearrangement mechanism tailored for sequential processing. In addition, we introduce a Mamba-based DETR head to improve multi-object representation. Evaluations on the nuScenes dataset demonstrate that MambaBEV-base achieves 51.7% NDS and an 42.7% mAP. Furthermore, evaluation within an end-to-end autonomous driving paradigm validates its effectiveness in motion forecasting and planning.These results highlight the potential of state-space models for improving global context understanding and large-object detection in autonomous driving perception systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。