arXiv:2504.21774cs.CVcs.LG2025-04中稿 · ITSC 2025被引 1

提出轻量级无人机协同感知框架,降低通信开销同时提升感知精度。

Is Intermediate Fusion All You Need for UAV-based Collaborative Perception?

  • 采用延迟-中间融合策略,交换紧凑检测结果并提升特征融合效率。
  • 在D435i-Dataset上实现93.6%的AP,通信量减少68%。
  • 适合资源受限的无人机群协同系统,尤其关注通信优化场景。

协同感知通过智能体间通信增强环境感知能力,是智能交通系统的重要方向。然而现有无人机协同方法未充分考虑无人机视角的独特性,导致通信开销大。为此,本文提出一种基于延迟-中间融合的高效协同感知框架LIF。核心思想是交换高信息量且紧凑的检测结果,并将融合阶段前移至特征表示层面。具体地,引入视觉引导的位置嵌入(VPE)和基于框的虚拟增强特征(BoBEV),有效融合多智能体互补信息;创新性地设计不确定性驱动的通信机制,通过评估不确定性筛选高质量共享区域。实验表明,LIF在保持极低通信带宽的前提下取得优异性能,验证了其有效性与实用性。代码与模型已开源:https://github.com/uestchjw/LIF。

原文摘要 · Abstract (English)

Collaborative perception enhances environmental awareness through inter-agent communication and is regarded as a promising solution to intelligent transportation systems. However, existing collaborative methods for Unmanned Aerial Vehicles (UAVs) overlook the unique characteristics of the UAV perspective, resulting in substantial communication overhead. To address this issue, we propose a novel communication-efficient collaborative perception framework based on late-intermediate fusion, dubbed LIF. The core concept is to exchange informative and compact detection results and shift the fusion stage to the feature representation level. In particular, we leverage vision-guided positional embedding (VPE) and box-based virtual augmented feature (BoBEV) to effectively integrate complementary information from various agents. Additionally, we innovatively introduce an uncertainty-driven communication mechanism that uses uncertainty evaluation to select high-quality and reliable shared areas. Experimental results demonstrate that our LIF achieves superior performance with minimal communication bandwidth, proving its effectiveness and practicality. Code and models are available at https://github.com/uestchjw/LIF.

无人机感知协同感知通信优化特征融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。