对比卫星影像与ImageNet预训练,发现前者优势有限。
Is Self-Supervised Pre-training on Satellite Imagery Better than ImageNet? A Systematic Study with Sentinel-2
- 在Sentinel-2影像上用SwAV和MAE做自监督预训练
- 六项下游任务中仅小幅提升性能,不如预期明显
- 适合关注遥感领域预训练成本效益的科研人员
自监督学习(SSL)在标注数据有限的情况下展现出了构建鲁棒模型的巨大潜力,对遥感(RS)任务尤为关键。普遍认为,在与下游任务同源的数据上进行预训练能带来最大收益,尤其优于ImageNet预训练(INP)。本文通过收集全球范围的光学哨兵-2(Sentinel-2)影像,构建了大型多样化的GeoNet数据集,并在该数据集与ImageNet上分别对SwAV和MAE进行预训练。在六项少样本下游任务上的评估显示,基于遥感数据的自监督预训练仅带来微弱性能提升,且在多种场景下仍保持竞争力。这表明遥感数据自监督预训练的优势可能被夸大,而数据整理与预训练带来的额外开销可能并不值得。
原文摘要 · Abstract (English)
Self-supervised learning (SSL) has demonstrated significant potential in pre-training robust models with limited labeled data, making it particularly valuable for remote sensing (RS) tasks. A common assumption is that pre-training on domain-aligned data provides maximal benefits on downstream tasks, particularly when compared to ImageNet-pretraining (INP). In this work, we investigate this assumption by collecting GeoNet, a large and diverse dataset of global optical Sentinel-2 imagery, and pre-training SwAV and MAE on both GeoNet and ImageNet. Evaluating these models on six downstream tasks in the few-shot setting reveals that SSL pre-training on RS data offers modest performance improvements over INP, and that it remains competitive in multiple scenarios. This indicates that the presumed benefits of SSL pre-training on RS data may be overstated, and the additional costs of data curation and pre-training could be unjustified.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。