用新型状态空间模型提升自动驾驶端到端感知效率与精度
GMF-Drive: Gated Mamba Fusion with Spatial-Aware BEV Representation for End-to-End Autonomous Driving
- 用几何增强的柱状表示替代传统激光雷达特征,保留三维结构信息
- 设计分层门控马比融合架构,实现线性复杂度长程依赖建模
- 在NAVSIM基准上超越现有方法,适合高要求自动驾驶系统
基于扩散模型的端到端自动驾驶正在引领新范式,但其性能受限于依赖变压器的特征融合机制。这类架构存在根本缺陷:二次计算复杂度限制了高分辨率特征使用,且缺乏空间先验,难以有效建模鸟瞰图(BEV)表征的内在结构。本文提出GMF-Drive(门控马比融合驱动),通过两项创新突破瓶颈。首先,将信息受限的直方图式激光雷达表示替换为编码形状描述符与统计特征的几何增强柱状格式,完整保留关键3D几何细节。其次,提出一种新型分层门控马比融合(GM-Fusion)架构,以高效、空间感知的状态空间模型(SSM)替代昂贵的变压器。核心的BEV-SSM利用方向序列与自适应融合机制,在线性复杂度下捕捉长程依赖,同时显式尊重驾驶场景的空间特性。在具有挑战性的NAVSIM基准上的大量实验表明,GMF-Drive达到新的最先进水平,显著优于DiffusionDrive。全面的消融实验验证各组件有效性,证明针对任务设计的SSM可在性能与效率上超越通用变压器。
原文摘要 · Abstract (English)
Diffusion-based models are redefining the state-of-the-art in end-to-end autonomous driving, yet their performance is increasingly hampered by a reliance on transformer-based fusion. These architectures face fundamental limitations: quadratic computational complexity restricts the use of high-resolution features, and a lack of spatial priors prevents them from effectively modeling the inherent structure of Bird's Eye View (BEV) representations. This paper introduces GMF-Drive (Gated Mamba Fusion for Driving), an end-to-end framework that overcomes these challenges through two principled innovations. First, we supersede the information-limited histogram-based LiDAR representation with a geometrically-augmented pillar format encoding shape descriptors and statistical features, preserving critical 3D geometric details. Second, we propose a novel hierarchical gated mamba fusion (GM-Fusion) architecture that substitutes an expensive transformer with a highly efficient, spatially-aware state-space model (SSM). Our core BEV-SSM leverages directional sequencing and adaptive fusion mechanisms to capture long-range dependencies with linear complexity, while explicitly respecting the unique spatial properties of the driving scene. Extensive experiments on the challenging NAVSIM benchmark demonstrate that GMF-Drive achieves a new state-of-the-art performance, significantly outperforming DiffusionDrive. Comprehensive ablation studies validate the efficacy of each component, demonstrating that task-specific SSMs can surpass a general-purpose transformer in both performance and efficiency for autonomous driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。