AdaBEV通过精炼与对比机制,提升多无人机目标检测的感知精度与效率。
Refine-and-Contrast: Adaptive Instance-Aware BEV Representations for Multi-UAV Collaborative Object Detection
- 基于框引导精炼和实例-背景对比学习,动态优化关键区域特征
- 在低分辨率输入下仍保持高精度,性能逼近理论上限
- 适合资源受限的无人机平台,计算开销极小
多无人机协同3D检测通过融合空中平台的多视角观测,实现更准确、鲁棒的感知,在覆盖范围和遮挡处理上具有显著优势,但对资源受限的无人机平台带来新的计算挑战。本文提出AdaBEV框架,通过精炼与对比范式学习自适应的实例感知鸟瞰图(BEV)表示。不同于现有方法对所有BEV网格同等处理,AdaBEV引入框引导精炼模块(BG-RM)和实例-背景对比学习(IBCL),以增强语义感知力与特征可区分性。BG-RM利用2D监督与空间分割,仅对前景实例关联的BEV网格进行精炼;IBCL则在BEV空间中通过对比学习强化前景与背景特征的分离。在Air-Co-Pred数据集上的大量实验表明,AdaBEV在不同模型规模下均实现优异的精度-计算权衡,于低分辨率下优于其他先进方法,并在保持低分辨率BEV输入和可忽略额外开销的前提下接近理论上限性能。
原文摘要 · Abstract (English)
Multi-UAV collaborative 3D detection enables accurate and robust perception by fusing multi-view observations from aerial platforms, offering significant advantages in coverage and occlusion handling, while posing new challenges for computation on resource-constrained UAV platforms. In this paper, we present AdaBEV, a novel framework that learns adaptive instance-aware BEV representations through a refine-and-contrast paradigm. Unlike existing methods that treat all BEV grids equally, AdaBEV introduces a Box-Guided Refinement Module (BG-RM) and an Instance-Background Contrastive Learning (IBCL) to enhance semantic awareness and feature discriminability. BG-RM refines only BEV grids associated with foreground instances using 2D supervision and spatial subdivision, while IBCL promotes stronger separation between foreground and background features via contrastive learning in BEV space. Extensive experiments on the Air-Co-Pred dataset demonstrate that AdaBEV achieves superior accuracy-computation trade-offs across model scales, outperforming other state-of-the-art methods at low resolutions and approaching upper bound performance while maintaining low-resolution BEV inputs and negligible overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。