arXiv:2503.13858cs.CVcs.LG2025-03ICLR被引 9

用状态空间模型高效学习鸟瞰图表示,提升多任务感知效率。

MamBEV: Enabling State Space Models to Learn Birds-Eye-View Representations

  • 基于Mamba的状态空间模型构建线性时空注意力,统一学习鸟瞰图特征。
  • 在多个3D感知任务中实现更优性能,输入规模扩展时效率远超基准模型。
  • 适合追求高效3D视觉感知的自动驾驶系统开发者使用。

3D视觉感知任务(如多摄像头图像中的3D检测)是自动驾驶与辅助系统的核心组成部分。然而,设计计算高效的算法仍是重大挑战。本文提出一种基于Mamba的框架MamBEV,利用线性时空状态空间模型(SSM)注意力机制,学习统一的鸟瞰图(BEV)表示,可支持多种3D感知任务,并显著提升计算与内存效率。此外,引入基于SSM的跨注意力机制,使BEV查询与图像特征实现有效交互,类似标准交叉注意力。大量实验表明,MamBEV在多种视觉感知指标上表现优异,尤其在输入规模扩展时展现出显著的效率优势,优于现有基准模型。

原文摘要 · Abstract (English)

3D visual perception tasks, such as 3D detection from multi-camera images, are essential components of autonomous driving and assistance systems. However, designing computationally efficient methods remains a significant challenge. In this paper, we propose a Mamba-based framework called MamBEV, which learns unified Bird's Eye View (BEV) representations using linear spatio-temporal SSM-based attention. This approach supports multiple 3D perception tasks with significantly improved computational and memory efficiency. Furthermore, we introduce SSM based cross-attention, analogous to standard cross attention, where BEV query representations can interact with relevant image features. Extensive experiments demonstrate MamBEV's promising performance across diverse visual perception metrics, highlighting its advantages in input scaling efficiency compared to existing benchmark models.

3D感知状态空间模型鸟瞰图自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。