用稀疏语义框传输,高效实现多车3D目标检测
Which2comm: An Efficient Collaborative Perception Framework for 3D Object Detection
- 用语义检测框传递物体级稀疏特征,减少通信量
- 在V2XSet和OPV2V上性能超越现有方法,通信成本更低
- 适合车联网、自动驾驶等低带宽场景应用
协同感知通过实时的多智能体信息交换,显著提升个体感知能力。然而实际场景中通信带宽有限,限制了数据传输量,导致协同感知系统性能下降,形成感知精度与通信开销的权衡。为此,我们提出Which2comm,一种基于物体级稀疏特征的多智能体3D目标检测框架。通过将物体语义信息融入3D检测框,引入语义检测框(SemDBs)。创新地在智能体间传输这些信息丰富的物体级稀疏特征,不仅大幅降低通信需求,还提升了3D目标检测性能。具体而言,构建全稀疏网络以从各智能体提取SemDBs;采用带有相对时间编码机制的时间融合方法,获得综合时空特征。在V2XSet和OPV2V数据集上的大量实验表明,Which2comm在感知性能与通信成本方面均持续优于现有先进方法,对真实世界延迟具有更强鲁棒性。结果表明,在多智能体协同3D目标检测中,仅传输物体级稀疏特征即可实现高精度且鲁棒的性能。
原文摘要 · Abstract (English)
Collaborative perception allows real-time inter-agent information exchange and thus offers invaluable opportunities to enhance the perception capabilities of individual agents. However, limited communication bandwidth in practical scenarios restricts the inter-agent data transmission volume, consequently resulting in performance declines in collaborative perception systems. This implies a trade-off between perception performance and communication cost. To address this issue, we propose Which2comm, a novel multi-agent 3D object detection framework leveraging object-level sparse features. By integrating semantic information of objects into 3D object detection boxes, we introduce semantic detection boxes (SemDBs). Innovatively transmitting these information-rich object-level sparse features among agents not only significantly reduces the demanding communication volume, but also improves 3D object detection performance. Specifically, a fully sparse network is constructed to extract SemDBs from individual agents; a temporal fusion approach with a relative temporal encoding mechanism is utilized to obtain the comprehensive spatiotemporal features. Extensive experiments on the V2XSet and OPV2V datasets demonstrate that Which2comm consistently outperforms other state-of-the-art methods on both perception performance and communication cost, exhibiting better robustness to real-world latency. These results present that for multi-agent collaborative 3D object detection, transmitting only object-level sparse features is sufficient to achieve high-precision and robust performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。