用自监督预训练和课程感知采样提升遥感SAR图像小目标检测性能
SAR Object Detection with Self-Supervised Pretraining and Curriculum-Aware Sampling
- 基于视觉变换器的自监督预训练,利用超25,700 km²无标注SAR数据
- 引入辅助二值分割任务,显著提升小目标检测精度
- 动态采样调度缓解类别不平衡,适合遥感小目标检测场景
星载合成孔径雷达(SAR)图像中的目标检测在城市监测与灾害响应中具有巨大潜力,但其固有复杂性及标注稀缺性带来了挑战。尤其小目标因分辨率低、噪声大而难检测。现有大规模标注SAR数据集匮乏,制约了深度学习模型发展。本文提出TRANSAR,一种基于视觉变换器的自监督端到端SAR目标检测模型,在超过25,700 km²地面面积的无标注SAR图像上进行掩码图像预训练。不同于传统检测范式,该方法引入辅助二值语义分割任务,用于后调优阶段分离关注目标(尤其是小目标)与背景。此外,为应对目标与图像尺寸比例失衡导致的类别不平衡问题,设计了基于课程学习与模型反馈的自适应采样调度器,动态调整训练时的目标类别分布。在多个基准SAR数据集上的大量实验表明,该方法优于传统监督模型(如DeepLabv3、UNet)及先进的自监督模型(如DPT、SegFormer、UperNet)。
原文摘要 · Abstract (English)
Object detection in satellite-borne Synthetic Aperture Radar (SAR) imagery holds immense potential in tasks such as urban monitoring and disaster response. However, the inherent complexities of SAR data and the scarcity of annotations present significant challenges in the advancement of object detection in this domain. Notably, the detection of small objects in satellite-borne SAR images poses a particularly intricate problem, because of the technology's relatively low spatial resolution and inherent noise. Furthermore, the lack of large labelled SAR datasets hinders the development of supervised deep learning-based object detection models. In this paper, we introduce TRANSAR, a novel self-supervised end-to-end vision transformer-based SAR object detection model that incorporates masked image pre-training on an unlabeled SAR image dataset that spans more than $25,700$ km\textsuperscript{2} ground area. Unlike traditional object detection formulation, our approach capitalises on auxiliary binary semantic segmentation, designed to segregate objects of interest during the post-tuning, especially the smaller ones, from the background. In addition, to address the innate class imbalance due to the disproportion of the object to the image size, we introduce an adaptive sampling scheduler that dynamically adjusts the target class distribution during training based on curriculum learning and model feedback. This approach allows us to outperform conventional supervised architecture such as DeepLabv3 or UNet, and state-of-the-art self-supervised learning-based arhitectures such as DPT, SegFormer or UperNet, as shown by extensive evaluations on benchmark SAR datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。