arXiv:2507.11267cs.CVcs.AI2025-07被引 1

YOLOatr提升热红外图像目标检测精度,实测达99.6%

YOLOatr : Deep Learning Based Automatic Target Detection and Localization in Thermal Infrared Imagery

  • 改进YOLOv5s结构,优化检测头与特征融合机制
  • 在DSIAC MWIR数据集上实现最高99.6%的检测准确率
  • 专为军事热成像场景设计,适合安防与侦察应用

在国防与监控领域,从热红外(TI)图像中实现自动目标检测(ATD)与识别(ATR)是极具挑战性的计算机视觉任务,相较于商用自动驾驶感知,受限于数据集稀缺、硬件限制、远距离导致的尺度不变性问题、战术车辆刻意遮挡、传感器分辨率低及目标结构信息缺失,以及天气、温差和昼夜变化影响,使类内差异增大、类间相似度升高,现有先进深度学习模型性能不足。本文提出一种改进的基于锚点的单阶段检测器YOLOatr,基于修改版YOLOv5s,优化检测头、颈部特征融合及定制增强策略。在DSIAC MWIR数据集上,针对相关与非相关测试协议进行评估,结果表明该模型在实时ATR任务中达到最高99.6%的SOTA性能。

原文摘要 · Abstract (English)

Automatic Target Detection (ATD) and Recognition (ATR) from Thermal Infrared (TI) imagery in the defense and surveillance domain is a challenging computer vision (CV) task in comparison to the commercial autonomous vehicle perception domain. Limited datasets, peculiar domain-specific and TI modality-specific challenges, i.e., limited hardware, scale invariance issues due to greater distances, deliberate occlusion by tactical vehicles, lower sensor resolution and resultant lack of structural information in targets, effects of weather, temperature, and time of day variations, and varying target to clutter ratios all result in increased intra-class variability and higher inter-class similarity, making accurate real-time ATR a challenging CV task. Resultantly, contemporary state-of-the-art (SOTA) deep learning architectures underperform in the ATR domain. We propose a modified anchor-based single-stage detector, called YOLOatr, based on a modified YOLOv5s, with optimal modifications to the detection heads, feature fusion in the neck, and a custom augmentation profile. We evaluate the performance of our proposed model on a comprehensive DSIAC MWIR dataset for real-time ATR over both correlated and decorrelated testing protocols. The results demonstrate that our proposed model achieves state-of-the-art ATR performance of up to 99.6%.

目标检测热红外YOLO军事应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。