arXiv:2511.19134cs.CV2025-11被引 1

融合可见光与红外图像,提升无人机小目标检测精度与速度。

MambaRefine-YOLO: A Dual-Modality Small Object Detector for UAV Imagery

  • 通过双门控机制动态融合可见光与红外信息,自适应平衡模态贡献。
  • 在双模态数据集上达到83.2%的mAP,较基线提升7.9%。
  • 结构轻量高效,适合实时无人机场景应用。

无人机影像中的小目标检测长期面临分辨率低和背景杂乱的挑战。虽然融合可见光(RGB)与红外(IR)数据具有潜力,但现有方法常在跨模态交互有效性与计算效率之间难以兼顾。本文提出MambaRefine-YOLO,核心包括:双门控互补马尔可夫融合模块(DGC-MFM),通过光照感知与差异感知门控机制自适应调节RGB与IR模态;以及分层特征聚合颈(HFAN),采用“先精炼后融合”策略增强多尺度特征。全面实验验证该双路径设计的有效性:在双模态DroneVehicle数据集上,全模型实现83.2%的mAP,较基线提升7.9%;在单模态VisDrone数据集上,仅使用HFAN的变体也取得显著性能提升,证明其通用性。本工作实现了精度与速度的优越平衡,适用于真实无人机应用场景。

原文摘要 · Abstract (English)

Small object detection in Unmanned Aerial Vehicle (UAV) imagery is a persistent challenge, hindered by low resolution and background clutter. While fusing RGB and infrared (IR) data offers a promising solution, existing methods often struggle with the trade-off between effective cross-modal interaction and computational efficiency. In this letter, we introduce MambaRefine-YOLO. Its core contributions are a Dual-Gated Complementary Mamba fusion module (DGC-MFM) that adaptively balances RGB and IR modalities through illumination-aware and difference-aware gating mechanisms, and a Hierarchical Feature Aggregation Neck (HFAN) that uses a ``refine-then-fuse'' strategy to enhance multi-scale features. Our comprehensive experiments validate this dual-pronged approach. On the dual-modality DroneVehicle dataset, the full model achieves a state-of-the-art mAP of 83.2%, an improvement of 7.9% over the baseline. On the single-modality VisDrone dataset, a variant using only the HFAN also shows significant gains, demonstrating its general applicability. Our work presents a superior balance between accuracy and speed, making it highly suitable for real-world UAV applications.

小目标检测无人机多模态融合YOLO

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。