arXiv:2409.07714cs.CVcs.MA2024-09被引 1

用状态空间模型提升多智能体协作感知效率,大幅降低计算通信开销。

CollaMamba: Efficient Collaborative Perception with Cross-Agent Spatial-Temporal State Space Model

  • 基于时空状态空间模型构建轻量级骨干网络,线性复杂度捕捉跨智能体依赖
  • 历史感知模块利用长时序上下文优化模糊特征,推理开销极低
  • 在多个数据集上精度更高,计算与通信开销分别降低71.9%和1/64,适合边缘部署

通过共享互补的感知信息,多智能体协作感知可深化对环境的理解。现有研究多采用CNN或Transformer在空间维度学习特征表示与融合,但在计算和通信资源受限下难以处理长程时空特征。全面建模大范围空间与长时间序列的依赖关系对提升特征质量至关重要。为此,我们提出一种资源高效的跨智能体时空状态空间模型(CollaMamba)。首先构建基于空间状态空间模型的骨干网络,能从单智能体与跨智能体视角有效捕获位置因果依赖,生成紧凑且全面的中间特征,同时保持线性复杂度。其次设计基于时间状态空间模型的历史感知特征增强模块,从扩展的历史帧中提取上下文线索以优化模糊特征,且开销极小。在多个数据集上的大量实验表明,CollaMamba优于当前最先进方法,在保持更高模型精度的同时,计算与通信开销分别降低最高达71.9%和1/64。该工作首次探索了Mamba在协作感知中的潜力。源代码将公开。

原文摘要 · Abstract (English)

By sharing complementary perceptual information, multi-agent collaborative perception fosters a deeper understanding of the environment. Recent studies on collaborative perception mostly utilize CNNs or Transformers to learn feature representation and fusion in the spatial dimension, which struggle to handle long-range spatial-temporal features under limited computing and communication resources. Holistically modeling the dependencies over extensive spatial areas and extended temporal frames is crucial to enhancing feature quality. To this end, we propose a resource efficient cross-agent spatial-temporal collaborative state space model (SSM), named CollaMamba. Initially, we construct a foundational backbone network based on spatial SSM. This backbone adeptly captures positional causal dependencies from both single-agent and cross-agent views, yielding compact and comprehensive intermediate features while maintaining linear complexity. Furthermore, we devise a history-aware feature boosting module based on temporal SSM, extracting contextual cues from extended historical frames to refine vague features while preserving low overhead. Extensive experiments across several datasets demonstrate that CollaMamba outperforms state-of-the-art methods, achieving higher model accuracy while reducing computational and communication overhead by up to 71.9% and 1/64, respectively. This work pioneers the exploration of the Mamba's potential in collaborative perception. The source code will be made available.

多智能体状态空间模型协作感知高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。