arXiv:2506.16737cs.CV2025-06被引 10

解决无人机多模态目标检测中的图像错位问题,提升检测精度。

Cross-modal Offset-guided Dynamic Alignment and Fusion for Weakly Aligned UAV Object Detection

论文配图:Cross-modal Offset-guided Dynamic Alignment and Fusion for Weakly Aligned UAV Object Detection
图 1 · 摘自论文原文
  • 用偏移引导的动态对齐模块,精准匹配可见光与红外图像特征。
  • 在DroneVehicle数据集上达到78.6%的mAP,显著优于现有方法。
  • 适合需要高鲁棒性无人机视觉检测的应用场景。

无人机目标检测在环境监测和城市安防中至关重要。为提升鲁棒性,近年研究通过融合可见光(RGB)与红外(IR)图像实现多模态检测。然而,由于无人机平台运动和成像异步,两模态常出现空间错位,导致语义不一致和特征融合冲突。现有方法通常分别处理,效果受限。本文提出跨模态偏移引导动态对齐与融合框架(CoDAF),包含两个新模块:偏移引导语义对齐(OSA)通过注意力估计空间偏移,利用共享语义空间引导可变形卷积实现精准对齐;动态注意力融合模块(DAFM)通过门控机制自适应平衡模态贡献,并结合空间-通道双重注意力优化融合特征。通过统一设计对齐与融合,CoDAF实现鲁棒的无人机目标检测。在标准基准测试中验证了有效性,其在DroneVehicle数据集上mAP达78.6%。

原文摘要 · Abstract (English)

Unmanned aerial vehicle (UAV) object detection plays a vital role in applications such as environmental monitoring and urban security. To improve robustness, recent studies have explored multimodal detection by fusing visible (RGB) and infrared (IR) imagery. However, due to UAV platform motion and asynchronous imaging, spatial misalignment frequently occurs between modalities, leading to weak alignment. This introduces two major challenges: semantic inconsistency at corresponding spatial locations and modality conflict during feature fusion. Existing methods often address these issues in isolation, limiting their effectiveness. In this paper, we propose Cross-modal Offset-guided Dynamic Alignment and Fusion (CoDAF), a unified framework that jointly tackles both challenges in weakly aligned UAV-based object detection. CoDAF comprises two novel modules: the Offset-guided Semantic Alignment (OSA), which estimates attention-based spatial offsets and uses deformable convolution guided by a shared semantic space to align features more precisely; and the Dynamic Attention-guided Fusion Module (DAFM), which adaptively balances modality contributions through gating and refines fused features via spatial-channel dual attention. By integrating alignment and fusion in a unified design, CoDAF enables robust UAV object detection. Experiments on standard benchmarks validate the effectiveness of our approach, with CoDAF achieving a mAP of 78.6% on the DroneVehicle dataset.

无人机检测多模态融合动态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。