arXiv:2409.10699cs.CV2024-09被引 13

用状态空间模型实现车载实时协同感知,效率更高。

CoMamba: Real-time Cooperative Perception Unlocked with State Space Models

  • 采用双向状态空间模型替代注意力机制,解决计算复杂度问题。
  • 在V2X数据集上检测精度优于现有方法,且保持实时处理能力。
  • 适合智能交通中大规模车辆协同感知场景使用。

协同感知系统在提升车联网安全与效率方面至关重要。尽管近年来车联万物(V2X)通信技术在自动驾驶中展现出显著效果,但核心挑战依然存在:如何高效融合不断扩展的联网设备(如车辆与基础设施)产生的高带宽特征。本文提出CoMamba,一种新型的协同3D检测框架,利用状态空间模型实现车载实时感知。相较于以往基于Transformer的先进模型,CoMamba采用双向状态空间模型,避免了注意力机制带来的二次复杂度瓶颈,具备更强可扩展性。在V2X/V2V数据集上的大量实验表明,CoMamba在保持实时处理能力的同时,性能显著优于现有方法。该框架不仅提升了目标检测精度,还大幅降低处理时间,为下一代智能交通网络中的协同感知系统提供了有力解决方案。

原文摘要 · Abstract (English)

Cooperative perception systems play a vital role in enhancing the safety and efficiency of vehicular autonomy. Although recent studies have highlighted the efficacy of vehicle-to-everything (V2X) communication techniques in autonomous driving, a significant challenge persists: how to efficiently integrate multiple high-bandwidth features across an expanding network of connected agents such as vehicles and infrastructure. In this paper, we introduce CoMamba, a novel cooperative 3D detection framework designed to leverage state-space models for real-time onboard vehicle perception. Compared to prior state-of-the-art transformer-based models, CoMamba enjoys being a more scalable 3D model using bidirectional state space models, bypassing the quadratic complexity pain-point of attention mechanisms. Through extensive experimentation on V2X/V2V datasets, CoMamba achieves superior performance compared to existing methods while maintaining real-time processing capabilities. The proposed framework not only enhances object detection accuracy but also significantly reduces processing time, making it a promising solution for next-generation cooperative perception systems in intelligent transportation networks.

协同感知状态空间模型3D检测实时系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。