用少量标注数据让SAM适应红外小目标检测,效果媲美全监督模型。
Scalpel-SAM: A Semi-Supervised Paradigm for Adapting SAM to Infrared Small Object Detection
- 设计分层MoE适配器,融合物理先验知识增强SAM
- 仅用10%标注数据训练出性能超全监督的教师模型
- 适合标注成本高、需快速部署的小目标检测场景
红外小目标检测因标注成本高,亟需半监督方法。现有SAM存在领域差异大、难以编码物理先验、结构复杂等问题。为此,我们设计了由四个白盒神经算子组成的分层MoE适配器,构建两阶段知识蒸馏与迁移范式:(1)先验引导知识蒸馏,利用MoE适配器和10%的全监督数据,将SAM蒸馏为专家教师模型(Scalpel-SAM);(2)部署导向知识迁移,用Scalpel-SAM生成伪标签,训练轻量高效的下游模型。实验表明,仅用极少标注,下游模型即可达到甚至超越全监督模型性能。据我们所知,这是首个系统性解决红外小目标检测数据稀缺问题、以SAM为教师模型的半监督范式。
原文摘要 · Abstract (English)
Infrared small object detection urgently requires semi-supervised paradigms due to the high cost of annotation. However, existing methods like SAM face significant challenges of domain gaps, inability of encoding physical priors, and inherent architectural complexity. To address this, we designed a Hierarchical MoE Adapter consisting of four white-box neural operators. Building upon this core component, we propose a two-stage paradigm for knowledge distillation and transfer: (1) Prior-Guided Knowledge Distillation, where we use our MoE adapter and 10% of available fully supervised data to distill SAM into an expert teacher (Scalpel-SAM); and (2) Deployment-Oriented Knowledge Transfer, where we use Scalpel-SAM to generate pseudo labels for training lightweight and efficient downstream models. Experiments demonstrate that with minimal annotations, our paradigm enables downstream models to achieve performance comparable to, or even surpassing, their fully supervised counterparts. To our knowledge, this is the first semi-supervised paradigm that systematically addresses the data scarcity issue in IR-SOT using SAM as the teacher model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。