用物理启发模块让视觉大模型适配红外小目标检测,兼顾单帧与视频场景。
SPIRIT: Adapting Vision Foundation Models for Unified Single- and Multi-Frame Infrared Small Target Detection
- 引入轻量级物理启发模块,增强红外弱信号特征。
- 在多个基准上超越现有方法,视频与单帧检测均表现优异。
- 适合需要统一处理红外小目标的安防与预警系统使用。
红外小目标检测(IRSTD)对监视和预警至关重要,应用场景涵盖单帧分析与视频追踪。理想方案应利用视觉基础模型(VFMs)缓解红外数据稀缺问题,并采用基于记忆-注意力的时序传播框架,统一支持单帧与多帧推理。然而,红外小目标辐射信号微弱、语义线索不足,与可见光图像差异显著,导致直接使用以语义为导向的VFMs及依赖外观的跨帧关联不可靠:层级特征聚合会淹没局部目标峰值,仅依赖外观的记忆注意力易引发虚假杂波关联。为此,我们提出SPIRIT,一种兼容视觉基础模型的统一框架,通过轻量级物理启发插件适配红外小目标检测。空间上,PIFR通过近似秩-稀疏分解,抑制结构化背景成分,增强类目标稀疏信号;时间上,PGMA将历史推导的软空间先验注入记忆交叉注意力,约束跨帧关联,在保证视频检测鲁棒性的同时,当无时序上下文时可自然退化为单帧推理。在多个IRSTD基准上的实验表明,该方法持续优于基于VFMs的基线并达到最先进性能。
原文摘要 · Abstract (English)
Infrared small target detection (IRSTD) is crucial for surveillance and early-warning, with deployments spanning both single-frame analysis and video-mode tracking. A practical solution should leverage vision foundation models (VFMs) to mitigate infrared data scarcity, while adopting a memory-attention-based temporal propagation framework that unifies single- and multi-frame inference. However, infrared small targets exhibit weak radiometric signals and limited semantic cues, which differ markedly from visible-spectrum imagery. This modality gap makes direct use of semantics-oriented VFMs and appearance-driven cross-frame association unreliable for IRSTD: hierarchical feature aggregation can submerge localized target peaks, and appearance-only memory attention becomes ambiguous, leading to spurious clutter associations. To address these challenges, we propose SPIRIT, a unified and VFM-compatible framework that adapts VFMs to IRSTD via lightweight physics-informed plug-ins. Spatially, PIFR refines features by approximating rank-sparsity decomposition to suppress structured background components and enhance sparse target-like signals. Temporally, PGMA injects history-derived soft spatial priors into memory cross-attention to constrain cross-frame association, enabling robust video detection while naturally reverting to single-frame inference when temporal context is absent. Experiments on multiple IRSTD benchmarks show consistent gains over VFM-based baselines and SOTA performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。