arXiv:2506.10425cs.CV2025-06中稿 · IEEE Transactions …被引 9

通过直接建模背景低秩结构,提升红外小目标检测精度与速度。

It's Not the Target, It's the Background: Rethinking Infrared Small Target Detection via Deep Patch-Free Low-Rank Representations

  • 不依赖分块处理,端到端学习图像域背景低秩表示
  • 在多个数据集上超越38种前沿方法,实时速度达82.34帧/秒
  • 对传感器噪声鲁棒,适合复杂背景下的实时检测场景

红外小目标检测(IRSTD)在复杂背景下长期面临信杂比低、目标形态多样、视觉特征缺失等挑战。现有深度学习方法虽致力于提取判别性表征,但小目标内在变化大、先验弱,常导致性能不稳定。本文提出新型端到端框架LRRNet,利用红外图像背景的低秩特性,借鉴杂波场景的物理可压缩性,采用压缩-重构-相减(CRS)范式,在图像域直接建模结构感知的低秩背景表示,无需分块处理或显式矩阵分解。据我们所知,这是首个以深度神经网络端到端学习低秩背景结构的工作。在多个公开数据集上的大量实验表明,LRRNet在检测精度、鲁棒性和计算效率上均优于38种先进方法。尤为突出的是,其平均速度达到82.34 FPS,实现实时处理。在具有挑战性的NoisySIRST数据集上的评估进一步验证了模型对传感器噪声的强鲁棒性。代码将在论文接受后公开。

原文摘要 · Abstract (English)

\textcolor{blue}{This is the pre-acceptance version, to read the final version please go to \href{https://ieeexplore.ieee.org/document/11156113}{IEEE Transactions on Geoscience and Remote Sensing on IEEE Xplore}.} Infrared small target detection (IRSTD) remains a long-standing challenge in complex backgrounds due to low signal-to-clutter ratios (SCR), diverse target morphologies, and the absence of distinctive visual cues. While recent deep learning approaches aim to learn discriminative representations, the intrinsic variability and weak priors of small targets often lead to unstable performance. In this paper, we propose a novel end-to-end IRSTD framework, termed LRRNet, which leverages the low-rank property of infrared image backgrounds. Inspired by the physical compressibility of cluttered scenes, our approach adopts a compression--reconstruction--subtraction (CRS) paradigm to directly model structure-aware low-rank background representations in the image domain, without relying on patch-based processing or explicit matrix decomposition. To the best of our knowledge, this is the first work to directly learn low-rank background structures using deep neural networks in an end-to-end manner. Extensive experiments on multiple public datasets demonstrate that LRRNet outperforms 38 state-of-the-art methods in terms of detection accuracy, robustness, and computational efficiency. Remarkably, it achieves real-time performance with an average speed of 82.34 FPS. Evaluations on the challenging NoisySIRST dataset further confirm the model's resilience to sensor noise. The source code will be made publicly available upon acceptance.

红外检测低秩表示实时处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。