提出新型注意力网络,提升红外小目标检测精度。
Selective Variable Convolution Meets Dynamic Content-Guided Attention for Infrared Small Target Detection
- 设计选择性可变卷积模块,扩大感受野并增强非局部特征
- 采用两阶段内容引导注意力机制,有效抑制误报
- 动态特征融合策略自适应整合上下文信息,适合复杂背景场景
红外小目标检测(IRSTD)旨在复杂背景下识别微小目标。传统卷积神经网络在处理该任务时,因卷积操作限制,难以充分提取小目标特征,导致关键信息丢失。为此,本文提出一种动态内容引导注意力多尺度特征聚合网络(DCGANet),遵循‘粗到精’注意力原则,实现高精度检测。首先,提出选择性可变卷积(SVC)模块,融合标准卷积、非规则可变形卷积与多速率空洞卷积优势,扩展感受野并强化非局部特征,有效提升目标与背景的区分能力。其次,核心组件为双阶段内容引导注意力模块:第一阶段聚焦特征图中的显著区域,第二阶段判断其是否为目标或背景干扰,通过保留关键响应抑制误报。此外,提出自适应动态特征融合(ADFF)模块替代静态特征级联,实现上下文特征的自适应融合,增强真目标判别能力。DCGANet在多个数据集上达到新基准。
原文摘要 · Abstract (English)
Infrared Small Target Detection (IRSTD) system aims to identify small targets in complex backgrounds. Due to the convolution operation in Convolutional Neural Networks (CNNs), applying traditional CNNs to IRSTD presents challenges, since the feature extraction of small targets is often insufficient, resulting in the loss of critical features. To address these issues, we propose a dynamic content-guided attention multiscale feature aggregation network (DCGANet), which adheres to the attention principle of 'coarse-to-fine' and achieves high detection accuracy. First, we propose a selective variable convolution (SVC) module that integrates the benefits of standard convolution, irregular deformable convolution, and multi-rate dilated convolution. This module is designed to expand the receptive field and enhance non-local features, thereby effectively improving the discrimination between targets and backgrounds. Second, the core component of DCGANet is a two-stage content-guided attention module. This module employs a two-stage attention mechanism to initially direct the network's focus to salient regions within the feature maps and subsequently determine whether these regions correspond to targets or background interference. By retaining the most significant responses, this mechanism effectively suppresses false alarms. Additionally, we propose an Adaptive Dynamic Feature Fusion (ADFF) module to substitute for static feature cascading. This dynamic feature fusion strategy enables DCGANet to adaptively integrate contextual features, thereby enhancing its ability to discriminate true targets from false alarms. DCGANet has achieved new benchmarks across multiple datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。