arXiv:2605.00886cs.CV2026-05

提出新网络提升红外小目标检测精度,解决背景干扰与细节丢失问题。

Selective Attention-Based Network for Robust Infrared Small Target Detection

  • 采用双路径结构融合局部细节与方向感知特征,增强目标表征能力。
  • 动态加权融合机制替代固定跳连,提升跨尺度特征利用效率。
  • 适合军事侦察、海上监视等需高精度小目标识别的场景。

红外小目标检测(IRSTD)在海事监控、军事搜救、预警系统和精确制导等领域至关重要,需在高度杂乱的红外背景下精准识别亮度低、仅占少数像素的目标。尽管深度学习取得进展,但红外小目标空间范围极小(常仅数像素)、信杂比低,易被复杂背景误判为虚假信号。现有编码器-解码器架构存在两大缺陷:早期卷积阶段的信息瓶颈削弱细粒度目标感知,静态跳跃连接缺乏动态适应性,难以区分真实目标与伪目标区域。为此,本文提出基于U-Net框架的Selective Attention Network(SANet),引入两个新组件:(1) 双路径语义感知模块(DSM),结合标准卷积保留局部空间细节与螺旋形卷积扩大方向敏感感受野,并通过卷积块注意力模块(CBAM)实现精细的空间-通道特征重校准;(2) 选择性注意力融合模块(SAFM),以空间自适应可学习权重机制替代传统静态跳跃连接,实现上下文感知的跨尺度特征融合。

原文摘要 · Abstract (English)

Infrared small target detection (IRSTD) plays a pivotal role in a broad spectrum of mission-critical applications, including maritime surveillance, military search and rescue, early warning systems, and precision-guided strikes, all of which demand the precise identification of dim, sub-pixel targets amid highly cluttered infrared backgrounds. Despite significant progress driven by deep learning methods, fundamental challenges persist: infrared small targets occupy extremely limited spatial extents (often only a few pixels), exhibit low signal-to-clutter ratios, and are easily confused with structurally complex backgrounds that frequently induce false alarms. Existing encoder-decoder architectures suffer from two key limitations - an information bottleneck in early convolutional stages that undermines fine-grained target perception, and static skip connections that lack the dynamic adaptability required to discriminate between genuine targets and pseudo-target regions. To address these challenges, we propose SANet, a Selective Attention-based Network built upon the classical U-Net framework and augmented with two novel components: (1) a \emph{Dual-path Semantic-aware Module} (DSM) that integrates standard convolutions for local spatial detail preservation with pinwheel-shaped convolutions for expanded, direction-sensitive receptive fields, followed by a Convolutional Block Attention Module (CBAM) for fine-grained spatial-channel feature recalibration; and (2) a \emph{Selective Attention Fusion Module} (SAFM) that replaces conventional static skip connections with a spatially adaptive, learnable weighting mechanism to perform context-aware, cross-scale feature fusion.

红外检测小目标注意力机制网络优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。