用视觉基础模型提升无源遥感目标检测的精度与稳定性
VFM-Guided Semi-Supervised Detection Transformer under Source-Free Constraints for Remote Sensing Object Detection
- 引入视觉基础模型引导伪标签生成,缓解噪声问题
- 在仅少量标注数据下实现优于现有方法的检测性能
- 适合遥感图像中目标密集、背景复杂的无源域适应场景
无监督域适应方法广泛用于弥合域间差异,但在真实遥感场景中,隐私和传输限制常导致无法获取源域数据,制约其应用。近年来,无源目标检测(SFOD)通过自训练范式不依赖源数据,成为有前景的替代方案。然而,面对遥感图像中密集目标和复杂背景,SFOD常因噪声伪标签导致训练崩溃。考虑到实际中可获取少量目标域标注数据,本文提出基于半监督框架的视觉基础模型引导检测变压器(VG-DETR)。该方法以“免费午餐”方式将视觉基础模型(VFM)融入训练流程,利用少量标注目标数据抑制伪标签噪声并增强特征提取能力。具体地,设计了基于VFM语义先验的伪标签挖掘策略,通过恢复低置信度输出中的潜在正确预测,提升伪标签质量和数量;同时提出双层级VFM引导对齐方法,在实例级和图像级分别通过对比学习细粒度原型与特征图相似性匹配,强化特征表示对域差异的鲁棒性。大量实验表明,VG-DETR在无源遥感目标检测任务中表现优异。
原文摘要 · Abstract (English)
Unsupervised domain adaptation methods have been widely explored to bridge domain gaps. However, in real-world remote-sensing scenarios, privacy and transmission constraints often preclude access to source domain data, which limits their practical applicability. Recently, Source-Free Object Detection (SFOD) has emerged as a promising alternative, aiming at cross-domain adaptation without relying on source data, primarily through a self-training paradigm. Despite its potential, SFOD frequently suffers from training collapse caused by noisy pseudo-labels, especially in remote sensing imagery with dense objects and complex backgrounds. Considering that limited target domain annotations are often feasible in practice, we propose a Vision foundation-Guided DEtection TRansformer (VG-DETR), built upon a semi-supervised framework for SFOD in remote sensing images. VG-DETR integrates a Vision Foundation Model (VFM) into the training pipeline in a "free lunch" manner, leveraging a small amount of labeled target data to mitigate pseudo-label noise while improving the detector's feature-extraction capability. Specifically, we introduce a VFM-guided pseudo-label mining strategy that leverages the VFM's semantic priors to further assess the reliability of the generated pseudo-labels. By recovering potentially correct predictions from low-confidence outputs, our strategy improves pseudo-label quality and quantity. In addition, a dual-level VFM-guided alignment method is proposed, which aligns detector features with VFM embeddings at both the instance and image levels. Through contrastive learning among fine-grained prototypes and similarity matching between feature maps, this dual-level alignment further enhances the robustness of feature representations against domain gaps. Extensive experiments demonstrate that VG-DETR achieves superior performance in source-free remote sensing detection tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。