arXiv:2603.06925cs.CV2026-03

融合可见光与红外图像,提升小目标检测精度与实时性

Small Target Detection Based on Mask-Enhanced Attention Fusion of Visible and Infrared Remote Sensing Images

  • 用可学习掩码和空间注意力实现像素级跨模态特征融合
  • 在VEDAI和DroneVehicle上分别达84.71%和74.0% mAP
  • 模型参数减少93.6%,计算量降低68%,适合部署

遥感图像中的目标通常尺寸小、纹理弱,易受复杂背景干扰,通用算法难以实现高精度检测。基于先前的ESM-YOLO,本文提出轻量级可见光-红外融合网络ESM-YOLO+。核心创新包括:(1) 提出掩码增强注意力融合(MEAF)模块,通过可学习空间掩码与空间注意力,在像素级融合可见光与红外特征,有效对齐跨模态信息,增强小目标表征,缓解跨模态错位与尺度异质性问题;(2) 引入训练时结构表示增强(SR),提供辅助监督以保留细粒度空间结构,提升特征判别力且不增加推理开销。在VEDAI和DroneVehicle数据集上的大量实验验证了其优越性:在VEDAI上达到84.71% mAP,DroneVehicle上达74.0% mAP,同时模型参数减少93.6%,计算量降低68.0%,证明ESM-YOLO+兼具高性能与实用性,为复杂遥感场景下的小目标检测提供了高效解决方案。

原文摘要 · Abstract (English)

Targets in remote sensing images are usually small, weakly textured, and easily disturbed by complex backgrounds, challenging high-precision detection with general algorithms. Building on our earlier ESM-YOLO, this work presents ESM-YOLO+ as a lightweight visible infrared fusion network. To enhance detection, ESM-YOLO+ includes two key innovations. (1) A Mask-Enhanced Attention Fusion (MEAF) module fuses features at the pixel level via learnable spatial masks and spatial attention, effectively aligning RGB and infrared features, enhancing small-target representation, and alleviating cross-modal misalignment and scale heterogeneity. (2) Training-time Structural Representation (SR) enhancement provides auxiliary supervision to preserve fine-grained spatial structures during training, boosting feature discriminability without extra inference cost. Extensive experiments on the VEDAI and DroneVehicle datasets validate ESM-YOLO+'s superiority. The model achieves 84.71\% mAP on VEDAI and 74.0\% mAP on DroneVehicle, while greatly reducing model complexity, with 93.6\% fewer parameters and 68.0\% lower GFLOPs than the baseline. These results confirm that ESM-YOLO+ integrates strong performance with practicality for real-time deployment, providing an effective solution for high-performance small-target detection in complex remote sensing scenes.

小目标检测多模态融合遥感图像轻量化模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。