arXiv:2504.14481cs.CV2025-04被引 1

通过形状先验提升SAM在红外小目标检测中的鲁棒性

LSP-ST: Ladder Shape-Biased Side-Tuning for Robust Infrared Small Target Detection

  • 引入全局结构先验替代手工边缘特征,增强模型对形状的感知
  • 仅用472万参数即达到顶尖性能,跨任务泛化能力强
  • 适合需要强结构感知的红外目标检测场景

将通用分割模型SAM微调用于红外小目标检测面临严重领域偏移问题。现有方法多依赖人工设计的先验,限制了泛化能力。本文发现基础模型存在纹理偏差,过度依赖局部纹理进行定位。为此提出阶梯式形状偏置侧调优(LSP-ST),引入形状感知归纳偏置,以全局结构先验融合边界与内部布局信息。设计形状增强大核注意力模块,以可微方式层次化隐式捕捉结构特征,无需任务特定人工引导。理论分析表明该注意力机制通过匹配滤波与反向传播提升结构感知学习能力。仅含472万可训练参数,LSP-ST在多个红外小目标检测基准上达到最先进性能,并在镜像、阴影、伪装物体检测等任务中展现强泛化能力,同时保持在显著性目标检测等纹理驱动任务上的稳定表现,证明形状偏置与纹理推理互补而非冲突。

原文摘要 · Abstract (English)

Fine-tuning the Segment Anything Model (SAM) for infrared small target detection poses significant challenges due to severe domain shifts. Existing adaptation methods often incorporate handcrafted priors to bridge this gap, yet such designs limit generalization and scalability. We identify a fundamental texture bias in foundation models, which overly depend on local texture cues for target localization. To address this, we propose Ladder Shape-Biased Side-Tuning (LSP-ST), a novel approach that introduces a shape-aware inductive bias to facilitate effective adaptation beyond texture cues. In contrast to prior work that injects explicit edge or contour features, LSP-ST models shape as a global structural prior, integrating both boundaries and internal layouts. We design a Shape-Enhanced Large-Kernel Attention Module to hierarchically and implicitly capture structural information in a fully differentiable manner, without task-specific handcrafted guidance. A theoretical analysis grounded in matched filtering and backpropagation reveals the mechanism by which the proposed attention improves structure-aware learning. With only 4.72M learnable parameters, LSP-ST achieves state-of-the-art performance on multiple infrared small target detection benchmarks. Furthermore, its strong generalization is validated across tasks such as mirror detection, shadow detection, and camouflaged object detection, while maintaining stable performance on texture-driven tasks like salient object detection, demonstrating that the introduced shape bias complements rather than competes with texture-based reasoning.

红外检测形状先验模型微调结构感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。