评估遥感基础模型跨区域作物分类能力,发现专用模型更优且少量标注即可生效。
On the Generalizability of Foundation Models for Crop Type Mapping
- 用多光谱卫星数据预训练的专用模型表现更好
- 仅需100张标注图像即达高准确率,900张可缓解类别不平衡
- 适合资源匮乏地区作物监测,关注地理泛化性
基于自监督学习预训练的基础模型在语言理解、文本生成和图像识别等任务中展现出强大的迁移能力。地球观测领域已开发出多个直接在多光谱卫星影像上预训练的遥感基础模型,应用于精准农业、野火与干旱监测、自然灾害响应等场景。然而,这些模型在新地理区域的泛化能力尚缺乏系统研究,存在地理偏差风险——即在数据丰富的发达国家训练的模型难以迁移到数据稀缺的发展中国家。本文在五大洲的五个作物分类数据集上评估了三种主流遥感基础模型:SSL4EO-S12、SatlasPretrain 和 ImageNet。结果表明,专为哨兵-2数据设计的预训练权重(如 SSL4EO-S12)优于通用预训练权重(如 ImageNet)。仅需100个标注样本即可达到高总体准确率,但要缓解类别不平衡问题并提升平均准确率,则需900个样本。
原文摘要 · Abstract (English)
Foundation models pre-trained using self-supervised learning have shown powerful transfer learning capabilities on various downstream tasks, including language understanding, text generation, and image recognition. The Earth observation (EO) field has produced several foundation models pre-trained directly on multispectral satellite imagery for applications like precision agriculture, wildfire and drought monitoring, and natural disaster response. However, few studies have investigated the ability of these models to generalize to new geographic locations, and potential concerns of geospatial bias -- models trained on data-rich developed nations not transferring well to data-scarce developing nations -- remain. We evaluate three popular EO foundation models, SSL4EO-S12, SatlasPretrain, and ImageNet, on five crop classification datasets across five continents. Results show that pre-trained weights designed explicitly for Sentinel-2, such as SSL4EO-S12, outperform general pre-trained weights like ImageNet. While only 100 labeled images are sufficient for achieving high overall accuracy, 900 images are required to mitigate class imbalance and improve average accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。