arXiv:2601.09228cs.CV2026-01

用语言提示分离红外图像中的物体与非物体特征,提升检测精度。

Disentangle Object and Non-object Infrared Features via Language Guidance

  • 引入文本监督,通过语义对齐引导物体特征提取
  • 在M³FD和FLIR数据集上分别达到83.7%和86.1% mAP
  • 适合需要高鲁棒性红外目标检测的场景

红外目标检测旨在复杂环境(如黑暗、雪天、雨天)中识别和定位物体,这些环境下可见光摄像头因光照不足而失效。然而,红外图像对比度低、边缘信息弱,难以提取具有判别性的物体特征。为此,我们提出一种新型视觉-语言表征学习范式用于红外目标检测。通过引入富含语义信息的文本监督,指导物体与非物体特征的解耦。具体而言,设计了语义特征对齐(SFA)模块,将物体特征与对应文本特征对齐;并开发了物体特征解耦(OFD)模块,通过最小化相关性实现文本对齐物体特征与非物体特征的解耦。最终,将解耦后的物体特征输入检测头,显著提升检测性能。大量实验表明,该方法在两个基准数据集上表现优异:M³FD上达到83.7% mAP,FLIR上达到86.1% mAP。代码将在论文接收后公开。

原文摘要 · Abstract (English)

Infrared object detection focuses on identifying and locating objects in complex environments (\eg, dark, snow, and rain) where visible imaging cameras are disabled by poor illumination. However, due to low contrast and weak edge information in infrared images, it is challenging to extract discriminative object features for robust detection. To deal with this issue, we propose a novel vision-language representation learning paradigm for infrared object detection. An additional textual supervision with rich semantic information is explored to guide the disentanglement of object and non-object features. Specifically, we propose a Semantic Feature Alignment (SFA) module to align the object features with the corresponding text features. Furthermore, we develop an Object Feature Disentanglement (OFD) module that disentangles text-aligned object features and non-object features by minimizing their correlation. Finally, the disentangled object features are entered into the detection head. In this manner, the detection performance can be remarkably enhanced via more discriminative and less noisy features. Extensive experimental results demonstrate that our approach achieves superior performance on two benchmarks: M\textsuperscript{3}FD (83.7\% mAP), FLIR (86.1\% mAP). Our code will be publicly available once the paper is accepted.

红外检测视觉语言特征解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。