用关键目标信息替代全图特征,大幅降低多车协同感知通信量。
CoCMT: Communication-Efficient Cross-Modal Transformer for Collaborative Perception
- 基于目标查询筛选关键信息,只传必要特征
- 在V2V4Real上仅需0.416Mb带宽,比顶尖方法少83倍
- 适合低带宽场景部署,检测精度还提升1.1%
多智能体协同感知通过共享传感信息提升各智能体的感知能力,有效应对传感器局限、遮挡和远距离感知等挑战。但现有系统传输如鸟瞰图(BEV)等中间特征,包含大量非关键信息,导致通信开销高。为此,我们提出CoCMT,一种基于目标查询的协同框架,通过选择性提取并传输核心特征来优化通信效率。其中引入高效查询变换器(EQFormer)融合多智能体目标查询,并采用协同深度监督增强阶段间正向强化,提升整体性能。在OPV2V和V2V4Real数据集上的实验表明,CoCMT优于当前最优方法,同时显著降低通信需求。在V2V4Real上,使用前50个目标查询的模型仅需0.416 Mb带宽,较SOTA减少83倍,且AP70提升1.1个百分点。这一效率突破使带宽受限环境下的实际协同感知部署成为可能,且不牺牲检测精度。
原文摘要 · Abstract (English)
Multi-agent collaborative perception enhances each agent perceptual capabilities by sharing sensing information to cooperatively perform robot perception tasks. This approach has proven effective in addressing challenges such as sensor deficiencies, occlusions, and long-range perception. However, existing representative collaborative perception systems transmit intermediate feature maps, such as bird-eye view (BEV) representations, which contain a significant amount of non-critical information, leading to high communication bandwidth requirements. To enhance communication efficiency while preserving perception capability, we introduce CoCMT, an object-query-based collaboration framework that optimizes communication bandwidth by selectively extracting and transmitting essential features. Within CoCMT, we introduce the Efficient Query Transformer (EQFormer) to effectively fuse multi-agent object queries and implement a synergistic deep supervision to enhance the positive reinforcement between stages, leading to improved overall performance. Experiments on OPV2V and V2V4Real datasets show CoCMT outperforms state-of-the-art methods while drastically reducing communication needs. On V2V4Real, our model (Top-50 object queries) requires only 0.416 Mb bandwidth, 83 times less than SOTA methods, while improving AP70 by 1.1 percent. This efficiency breakthrough enables practical collaborative perception deployment in bandwidth-constrained environments without sacrificing detection accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。