用语言提示提升红外小目标检测,显著改善精度。
Leveraging Language Prior for Infrared Small Target Detection
- 引入语言先验引导注意力,融合文本与图像信息
- 在两个数据集上检测精度提升超67%,尤其大幅降低误报率
- 构建首个多模态红外小目标数据集,推动领域发展
红外小目标检测(IRSTD)在模糊背景中识别微小目标,对多种应用至关重要。由于目标尺寸小且分布稀疏,该任务极具挑战性。现有方法和数据集仅依赖图像模态,限制了性能提升。本文提出一种新型多模态IRSTD框架,利用语言先验引导检测过程。通过GPT-4生成描述目标位置的文本,并经精心设计提示工程,提取语言引导注意力权重,增强模型感知能力。针对缺乏多模态红外数据集的问题,我们构建了包含图像与文本的LangIR数据集,扩展自IRSTD-1k与NUDT-SIRST。大量实验与消融研究验证了有效性:在NUAA-SIRST子集上,IoU、nIoU、Pd、Fa分别提升9.74%、13.02%、1.25%、67.87%;在IRSTD-1k子集上,相应提升为4.41%、2.04%、2.01%、113.43%。
原文摘要 · Abstract (English)
IRSTD (InfraRed Small Target Detection) detects small targets in infrared blurry backgrounds and is essential for various applications. The detection task is challenging due to the small size of the targets and their sparse distribution in infrared small target datasets. Although existing IRSTD methods and datasets have led to significant advancements, they are limited by their reliance solely on the image modality. Recent advances in deep learning and large vision-language models have shown remarkable performance in various visual recognition tasks. In this work, we propose a novel multimodal IRSTD framework that incorporates language priors to guide small target detection. We leverage language-guided attention weights derived from the language prior to enhance the model's ability for IRSTD, presenting a novel approach that combines textual information with image data to improve IRSTD capabilities. Utilizing the state-of-the-art GPT-4 vision model, we generate text descriptions that provide the locations of small targets in infrared images, employing careful prompt engineering to ensure improved accuracy. Due to the absence of multimodal IR datasets, existing IRSTD methods rely solely on image data. To address this shortcoming, we have curated a multimodal infrared dataset that includes both image and text modalities for small target detection, expanding upon the popular IRSTD-1k and NUDT-SIRST datasets. We validate the effectiveness of our approach through extensive experiments and comprehensive ablation studies. The results demonstrate significant improvements over the state-of-the-art method, with relative percentage differences of 9.74%, 13.02%, 1.25%, and 67.87% in IoU, nIoU, Pd, and Fa on the NUAA-SIRST subset, and 4.41%, 2.04%, 2.01%, and 113.43% on the IRSTD-1k subset of the LangIR dataset, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。