融合红外与可见光图像,用注意力机制提升小目标检测精度
Enhanced Small Target Detection via Multi-Modal Fusion and Attention Mechanisms: A YOLOv5 Approach
- 用特征点匹配对齐红外与可见光图像,实现多模态数据融合
- 在Visdrone和anti-UAV数据集上小目标检测精度显著提升
- 适合军事侦察、无人机反制等复杂环境下的实时检测场景
随着信息技术快速发展,现代战争越来越依赖情报,小目标检测在军事应用中至关重要。由于复杂环境下干扰增多,高效实时检测小目标面临挑战。为此,本文提出一种基于多模态图像融合与注意力机制的小目标检测方法,采用YOLOv5框架,结合红外与可见光数据,并引入卷积注意力模块以增强检测性能。首先通过特征点匹配进行多模态数据集配准,确保网络训练准确。融合红外与可见光特征并结合注意力机制后,模型在检测精度和鲁棒性上均得到提升。在anti-UAV与Visdrone数据集上的实验结果表明,该方法在小目标和暗弱目标检测方面表现优异,验证了其有效性和实用性。
原文摘要 · Abstract (English)
With the rapid development of information technology, modern warfare increasingly relies on intelligence, making small target detection critical in military applications. The growing demand for efficient, real-time detection has created challenges in identifying small targets in complex environments due to interference. To address this, we propose a small target detection method based on multi-modal image fusion and attention mechanisms. This method leverages YOLOv5, integrating infrared and visible light data along with a convolutional attention module to enhance detection performance. The process begins with multi-modal dataset registration using feature point matching, ensuring accurate network training. By combining infrared and visible light features with attention mechanisms, the model improves detection accuracy and robustness. Experimental results on anti-UAV and Visdrone datasets demonstrate the effectiveness and practicality of our approach, achieving superior detection results for small and dim targets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。