提出TDCNet,用时间差卷积提升红外小目标检测精度
Spatio-Temporal Context Learning with Temporal Difference Convolution for Moving Infrared Small Target Detection
- 设计时间差卷积模块,融合时序差异与3D卷积提取多尺度运动特征
- 在IRSTD-UAV和公开数据集上达到当前最优性能,漏检率显著降低
- 适合无人机红外目标检测、复杂背景下的弱小目标识别场景
移动红外小目标检测(IRSTD)在无人机监视与搜救系统中至关重要,但因目标特征微弱且背景干扰复杂而极具挑战。准确的时空特征建模对检测效果至关重要,传统方法依赖时序差分或三维卷积:前者虽能显式利用运动信息,但空间特征提取能力有限;后者虽能有效表示时空特征,却缺乏对时序动态的显式感知。本文提出新型移动红外小目标检测网络TDCNet,引入可重参数化的时间差卷积(TDC)模块,包含三个并行的TDC块,分别捕捉不同时间范围的上下文依赖。每个块将时间差分与3D卷积融合为统一的时空卷积表示,有效提取多尺度运动上下文特征,同时抑制复杂背景中的伪运动杂波。此外,提出基于TDC的时空注意力机制,实现主干网络与平行3D主干间特征的跨注意力交互,建模全局语义依赖以优化当前帧特征。在IRSTD-UAV及多个公开红外数据集上的大量实验表明,TDCNet在移动目标检测任务中达到最新水平,显著提升检测性能。
原文摘要 · Abstract (English)
Moving infrared small target detection (IRSTD) plays a critical role in practical applications, such as surveillance of unmanned aerial vehicles (UAVs) and UAV-based search system. Moving IRSTD still remains highly challenging due to weak target features and complex background interference. Accurate spatio-temporal feature modeling is crucial for moving target detection, typically achieved through either temporal differences or spatio-temporal (3D) convolutions. Temporal difference can explicitly leverage motion cues but exhibits limited capability in extracting spatial features, whereas 3D convolution effectively represents spatio-temporal features yet lacks explicit awareness of motion dynamics along the temporal dimension. In this paper, we propose a novel moving IRSTD network (TDCNet), which effectively extracts and enhances spatio-temporal features for accurate target detection. Specifically, we introduce a novel temporal difference convolution (TDC) re-parameterization module that comprises three parallel TDC blocks designed to capture contextual dependencies across different temporal ranges. Each TDC block fuses temporal difference and 3D convolution into a unified spatio-temporal convolution representation. This re-parameterized module can effectively capture multi-scale motion contextual features while suppressing pseudo-motion clutter in complex backgrounds, significantly improving detection performance. Moreover, we propose a TDC-guided spatio-temporal attention mechanism that performs cross-attention between the spatio-temporal features from the TDC-based backbone and a parallel 3D backbone. This mechanism models their global semantic dependencies to refine the current frame's features. Extensive experiments on IRSTD-UAV and public infrared datasets demonstrate that our TDCNet achieves state-of-the-art detection performance in moving target detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。