利用目标随时间逐渐显现的特性,提升红外视频中小目标检测精度。
Temporal-Emerged Prompting for Segment Anything in Multiframe Infrared Small Target Detection

- 通过建模全局运动与局部运动差异定位潜在目标
- 在低信噪比下仍能实现精准分割,复杂背景表现优异
- 适合红外小目标检测任务,无需交互式标注
在低信噪比红外序列中精确定位和分割小目标仍是难题。由于目标在单帧中常与背景难以区分,现有方法即使使用先进基础模型和强大跨帧关联机制,仍难以检测。受目标随时间逐步从背景中浮现并变得可辨认的启发,我们提出时序涌现提示框架(TEP-SAM),旨在显式利用这种时序涌现特征来调制和提示分割一切模型(SAM)。TEP-SAM通过联合建模全局运动模式与局部运动偏差来定位潜在目标,并借助运动差异增强目标区域特征,生成用于SAM的时序涌现提示,实现非交互式分割。通过融合大规模语义预训练与任务特定时序建模,TEP-SAM有效适配了SAM在多帧红外小目标检测任务中的应用。大量实验表明,该方法在严重低信噪比条件及复杂动态背景中均表现出色。
原文摘要 · Abstract (English)
Accurately localizing and segmenting small targets in low signal-to-noise ratio (SNR) infrared sequences remains a challenging task. Since targets are often indistinguishable from the background in individual frames, existing methods, even when equipped with advanced foundation model and powerful inter-frame association mechanisms, still fail to detect them. Motivated by the observation that targets tend to emerge gradually from the background over time and become distinguishable, we propose Temporal-Emerged Prompting for Segment Anything Model (TEP-SAM), a principled framework designed to explicitly exploit such temporal-emerged cues to modulate and prompt SAM. TEP-SAM operates by jointly modeling global motion patterns and local motion deviations to locate potential targets. It further enhances target region features by leveraging motion discrepancy, thereby generating temporal-emerged cues for SAM and enabling non-interactive segmentation. By bridging large-scale semantic pretraining with task-specific temporal modeling, TEP-SAM effectively adapts SAM to the challenging multiframe infrared small target detection task. Extensive experiments demonstrate the effectiveness of our approach, particularly under severely low-SNR conditions and in complex dynamic background.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。