arXiv:2511.20319cs.CV2025-11

动态适配红外图像状态,提升小目标检测精度

IrisNet: Infrared Image Status Awareness Meta Decoder for Infrared Small Targets Detection

  • 用图像-解码器变换器建立输入与解码参数的动态映射
  • 在三个数据集上达到当前最好性能,显著提升小目标检出率
  • 适合复杂场景下红外小目标检测,尤其夜间/多域应用

红外小目标检测因信噪比低、背景复杂及目标特征不明显而面临挑战。尽管基于深度学习的编码器-解码器框架取得进展,但其静态模式学习在不同场景(如昼夜变化、天空/海面/陆地域)下易出现模式漂移,影响鲁棒性。为此,本文提出IrisNet,一种新型元学习框架,可动态适应输入红外图像的状态。通过图像-解码器变换器建立红外图像特征与完整解码器参数间的动态映射。具体而言,将参数化解码器表示为保留层级相关性的二维张量,利用自注意力建模层间依赖,并通过交叉注意力生成自适应解码模式。为进一步增强红外图像感知能力,融合高频成分以补充目标位置与场景边缘信息。在NUDT-SIRST、NUAA-SIRST和IRSTD-1K数据集上的实验表明,IrisNet表现优异,达到当前最优水平。

原文摘要 · Abstract (English)

Infrared Small Target Detection (IRSTD) faces significant challenges due to low signal-to-noise ratios, complex backgrounds, and the absence of discernible target features. While deep learning-based encoder-decoder frameworks have advanced the field, their static pattern learning suffers from pattern drift across diverse scenarios (\emph{e.g.}, day/night variations, sky/maritime/ground domains), limiting robustness. To address this, we propose IrisNet, a novel meta-learned framework that dynamically adapts detection strategies to the input infrared image status. Our approach establishes a dynamic mapping between infrared image features and entire decoder parameters via an image-to-decoder transformer. More concretely, we represent the parameterized decoder as a structured 2D tensor preserving hierarchical layer correlations and enable the transformer to model inter-layer dependencies through self-attention while generating adaptive decoding patterns via cross-attention. To further enhance the perception ability of infrared images, we integrate high-frequency components to supplement target-position and scene-edge information. Experiments on NUDT-SIRST, NUAA-SIRST, and IRSTD-1K datasets demonstrate the superiority of our IrisNet, achieving state-of-the-art performance.

红外检测小目标元学习动态解码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。