卫星任务无需专用预训练,通用模型已足够好。
Do Satellite Tasks Need Special Pretraining?
- 用低分辨率图像测试卫星模型泛化能力,评估下游任务表现。
- 在百万级遥感数据集上训练iBOT,但未超越通用视觉模型。
- 在小规模下,专用预训练无明显优势,适合资源有限的研究者。
基础模型已在多种模态中推动机器学习发展,包括图像。近期多个团队训练了针对遥感应用的专用基础模型。这一研究方向源于遥感图像的独特特征、特定应用场景以及对鲁棒性的特殊需求。本文系统挑战了‘专用基础模型优于通用视觉基础模型’的观点,至少在小规模场景下如此。首先,设计了一个简单基准,评估遥感模型在低分辨率图像上的泛化能力,涵盖两个下游任务。其次,在百万级遥感图像数据集MillionAID上,对iBOT(一种自监督视觉编码器)进行训练,并针对遥感特性做了若干调整。结果显示,这些预训练模型在ViT-B规模下,未能持续优于通用视觉基线模型。
原文摘要 · Abstract (English)
Foundation models have advanced machine learning across various modalities, including images. Recently multiple teams trained foundation models specialized for remote sensing applications. This line of research is motivated by the distinct characteristics of remote sensing imagery, specific applications and types of robustness useful for satellite image analysis. In this work we systematically challenge the idea that specific foundation models are more useful than general-purpose vision foundation models, at least in the small scale. First, we design a simple benchmark that measures generalization of remote sensing models towards images with lower resolution for two downstream tasks. Second, we train iBOT, a self-supervised vision encoder, on MillionAID, an ImageNet-scale satellite imagery dataset, with several modifications specific to remote sensing. We show that none of those pretrained models bring consistent improvements upon general-purpose baselines at the ViT-B scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。