arXiv:2511.19306cs.CV2025-11

用双粒度语义提示提升红外小目标检测精度,无需人工标注

Dual-Granularity Semantic Prompting for Language Guidance Infrared Small Target Detection

  • 设计粗粒度与细粒度结合的文本提示,实现无标注语言引导
  • 在三个基准数据集上达到当前最优检测性能
  • 适合关注红外目标检测与多模态融合的研究者

红外小目标检测因特征表达有限和背景干扰严重而面临挑战,导致性能不佳。尽管近期受CLIP启发的方法尝试利用文本指导进行检测,但仍受限于文本描述不准确及依赖人工标注。为此,我们提出DGSPNet,一种端到端的语言提示驱动框架。该方法融合双粒度语义提示:粗粒度文本先验(如‘红外图像’、‘小目标’)和通过图像空间内视觉到文本映射生成的细粒度个性化语义描述。这一设计不仅有助于学习细粒度语义信息,还能在推理时自然利用语言提示,无需任何标注。通过充分挖掘文本描述的精准性与简洁性,我们进一步引入文本引导通道注意力(TGCA)和文本引导空间注意力(TGSA)机制,增强模型在低层与高层特征空间中对潜在目标的敏感度。大量实验表明,该方法显著提升检测精度,在三个基准数据集上达到当前最优表现。

原文摘要 · Abstract (English)

Infrared small target detection remains challenging due to limited feature representation and severe background interference, resulting in sub-optimal performance. While recent CLIP-inspired methods attempt to leverage textual guidance for detection, they are hindered by inaccurate text descriptions and reliance on manual annotations. To overcome these limitations, we propose DGSPNet, an end-to-end language prompt-driven framework. Our approach integrates dual-granularity semantic prompts: coarse-grained textual priors (e.g., 'infrared image', 'small target') and fine-grained personalized semantic descriptions derived through visual-to-textual mapping within the image space. This design not only facilitates learning fine-grained semantic information but also can inherently leverage language prompts during inference without relying on any annotation requirements. By fully leveraging the precision and conciseness of text descriptions, we further introduce a text-guide channel attention (TGCA) mechanism and text-guide spatial attention (TGSA) mechanism that enhances the model's sensitivity to potential targets across both low- and high-level feature spaces. Extensive experiments demonstrate that our method significantly improves detection accuracy and achieves state-of-the-art performance on three benchmark datasets.

红外检测语义提示多模态注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。