用目标数量提示实现弱监督,提升红外小目标检测精度
Weakly-supervised Contrastive Learning with Quantity Prompts for Moving Infrared Small Target Detection
- 基于SAM模型挖掘潜在目标,结合多帧能量累积
- 对比学习增强伪标签可靠性,性能达SOTA的90%以上
- 适合标注成本高的红外视频小目标检测场景
与通用目标检测不同,移动红外小目标检测因目标尺寸微小、背景对比度弱而面临巨大挑战。现有方法多为全监督,严重依赖大量人工标注,而视频序列的人工标注在低质量红外图像上尤为耗时费力。受通用目标检测启发,本文首次探索弱监督策略以减少标注需求。提出一种新的弱监督对比学习框架WeCoL,仅需简单的目标数量提示即可训练。具体地,基于预训练的Segment Anything Model(SAM),设计了结合目标激活图与多帧能量累积的潜在目标挖掘策略;采用对比学习在特征子空间中计算正负样本相似性,进一步提升伪标签可靠性;同时提出长短时运动感知学习机制,同时建模小目标的局部运动模式与全局运动轨迹。在两个公开数据集DAUB和ITSDT-15K上的大量实验表明,该弱监督方案常优于早期全监督方法,性能可达当前最先进全监督方法的90%以上。
原文摘要 · Abstract (English)
Different from general object detection, moving infrared small target detection faces huge challenges due to tiny target size and weak background contrast.Currently, most existing methods are fully-supervised, heavily relying on a large number of manual target-wise annotations. However, manually annotating video sequences is often expensive and time-consuming, especially for low-quality infrared frame images. Inspired by general object detection, non-fully supervised strategies ($e.g.$, weakly supervised) are believed to be potential in reducing annotation requirements. To break through traditional fully-supervised frameworks, as the first exploration work, this paper proposes a new weakly-supervised contrastive learning (WeCoL) scheme, only requires simple target quantity prompts during model training.Specifically, in our scheme, based on the pretrained segment anything model (SAM), a potential target mining strategy is designed to integrate target activation maps and multi-frame energy accumulation.Besides, contrastive learning is adopted to further improve the reliability of pseudo-labels, by calculating the similarity between positive and negative samples in feature subspace.Moreover, we propose a long-short term motion-aware learning scheme to simultaneously model the local motion patterns and global motion trajectory of small targets.The extensive experiments on two public datasets (DAUB and ITSDT-15K) verify that our weakly-supervised scheme could often outperform early fully-supervised methods. Even, its performance could reach over 90\% of state-of-the-art (SOTA) fully-supervised ones.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。