arXiv:2505.05397cs.CV2025-05被引 3

用混合状态空间模型提升路侧点云的局部全局感知能力

PillarMamba: Learning Local-Global Context for Roadside Point Cloud via Hybrid State Space Model

  • 设计跨阶段状态空间组,融合多层级特征提升表达能力
  • 在DAIR-V2X-I数据集上超越现有方法,显著提升检测精度
  • 适合关注路侧智能感知与高效点云处理的研究者

面向智能交通系统(ITS)与车联万物(V2X)任务,路侧感知近年来受到越来越多关注,因其可扩展联网车辆的感知范围并提升交通安全。然而,针对路侧点云的3D目标检测尚未得到充分探索。网络的感受野与场景上下文利用能力在很大程度上决定了检测性能。近期基于状态空间模型(SSM)的Mamba因其高效的全局感受野,对传统卷积与Transformer架构构成挑战。本文将Mamba引入基于柱状体的路侧点云感知,提出一种基于跨阶段状态空间组(CSG)的框架,命名为PillarMamba,通过跨阶段特征融合增强网络表达力并实现高效计算。然而,由于扫描方向限制,状态空间模型存在局部连接断裂与历史关系遗忘问题。为此,本文提出混合状态空间块(HSB),通过局部卷积增强邻域连接,利用残差注意力保留历史记忆,以获取路侧点云的局部-全局上下文。所提方法在主流大规模路侧基准数据集DAIR-V2X-I上优于当前最先进方法,代码即将开源。

原文摘要 · Abstract (English)

Serving the Intelligent Transport System (ITS) and Vehicle-to-Everything (V2X) tasks, roadside perception has received increasing attention in recent years, as it can extend the perception range of connected vehicles and improve traffic safety. However, roadside point cloud oriented 3D object detection has not been effectively explored. To some extent, the key to the performance of a point cloud detector lies in the receptive field of the network and the ability to effectively utilize the scene context. The recent emergence of Mamba, based on State Space Model (SSM), has shaken up the traditional convolution and transformers that have long been the foundational building blocks, due to its efficient global receptive field. In this work, we introduce Mamba to pillar-based roadside point cloud perception and propose a framework based on Cross-stage State-space Group (CSG), called PillarMamba. It enhances the expressiveness of the network and achieves efficient computation through cross-stage feature fusion. However, due to the limitations of scan directions, state space model faces local connection disrupted and historical relationship forgotten. To address this, we propose the Hybrid State-space Block (HSB) to obtain the local-global context of roadside point cloud. Specifically, it enhances neighborhood connections through local convolution and preserves historical memory through residual attention. The proposed method outperforms the state-of-the-art methods on the popular large scale roadside benchmark: DAIR-V2X-I. The code will be released soon.

点云检测状态空间模型路侧感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。