arXiv:2503.07249cs.CV2025-03ICCV被引 9

用语义文本提升复杂场景红外小目标检测精度

Text-IRSTD: Leveraging Semantic Text to Promote Infrared Small Target Detection in Complex Scenes

  • 引入模糊语义文本提示,解决目标类别不明确问题
  • 提出跨模态交互解码器,实现图文信息深度融合
  • 构建新数据集FZDT,验证方法在未见场景的泛化能力

红外小目标检测是计算机视觉中的热点与挑战任务。现有方法多依赖目标的视觉特征,难以应对复杂多变的检测场景。主要原因在于红外小目标自身图像信息有限,仅靠视觉特征难以区分目标与干扰,导致性能下降。为此,本文提出Text-IRSTD,首次将语义文本引入红外小目标检测,拓展经典方法为文本引导的红外小目标检测框架。一方面设计模糊语义文本提示以适应不确定的目标类别;另一方面提出渐进式跨模态语义交互解码器(PCSID),促进文本与图像的信息融合。此外,构建包含2,755张不同场景红外图像的新基准数据集FZDT,附带模糊语义文本标注。大量实验表明,所提方法在检测性能和目标轮廓恢复上均优于当前最优方法,且在未见场景中表现出强泛化能力。论文接受后将公开数据集与代码。

原文摘要 · Abstract (English)

Infrared small target detection is currently a hot and challenging task in computer vision. Existing methods usually focus on mining visual features of targets, which struggles to cope with complex and diverse detection scenarios. The main reason is that infrared small targets have limited image information on their own, thus relying only on visual features fails to discriminate targets and interferences, leading to lower detection performance. To address this issue, we introduce a novel approach leveraging semantic text to guide infrared small target detection, called Text-IRSTD. It innovatively expands classical IRSTD to text-guided IRSTD, providing a new research idea. On the one hand, we devise a novel fuzzy semantic text prompt to accommodate ambiguous target categories. On the other hand, we propose a progressive cross-modal semantic interaction decoder (PCSID) to facilitate information fusion between texts and images. In addition, we construct a new benchmark consisting of 2,755 infrared images of different scenarios with fuzzy semantic textual annotations, called FZDT. Extensive experimental results demonstrate that our method achieves better detection performance and target contour recovery than the state-of-the-art methods. Moreover, proposed Text-IRSTD shows strong generalization and wide application prospects in unseen detection scenarios. The dataset and code will be publicly released after acceptance of this paper.

红外检测文本引导跨模态小目标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。