提出新预训练策略,提升遥感图像分割模型泛化能力
A generalised pre-training strategy for deep learning networks in semantic segmentation of remotely sensed images
- 通过引导模型避开特定领域特征,改进ImageNet预训练
- 在4个不同场景数据集上均达最优表现,最高mIoU达84.22%
- 适合想提升遥感图像分割通用性的研究人员使用
在遥感图像语义分割中,深度学习模型通常在ImageNet等大规模图像数据库上预训练,再在特定领域数据集上微调。然而,ImageNet与遥感图像之间存在显著领域差距(如场景和模态差异),导致微调后性能受限。为此,研究者尝试构建大规模领域专用数据集用于预训练,但此类数据集建立困难且泛化能力有限。本文提出一种新颖而简单的预训练策略,旨在预训练阶段引导模型避免学习特定领域特征,从而提升预训练模型的泛化能力。为验证效果,模型在ImageNet上预训练后,在包含iSAID、MFNet、PST900和Potsdam在内的四个具有不同场景与模态的语义分割数据集上进行微调。实验结果表明,该策略在所有数据集上均达到当前最优性能:iSAID为67.4% mIoU,MFNet为56.9% mIoU,PST900为84.22% mIoU,Potsdam为91.88% mF1。本研究为构建适用于计算机视觉与遥感的统一基础模型奠定基础。
原文摘要 · Abstract (English)
In the segmentation of remotely sensed images, deep learning models are typically pre-trained using large image databases like ImageNet before fine-tuned on domain-specific datasets. However, the performance of these fine-tuned models is often hindered by the large domain gaps (i.e., differences in scenes and modalities) between ImageNet's images and remotely sensed images being processed. Therefore, many researchers have undertaken efforts to establish large-scale domain-specific image datasets for pre-training, aiming to enhance model performance. However, establishing such datasets is often challenging, requiring significant effort, and these datasets often exhibit limited generaliza-bility to other application scenarios. To address these issues, this study introduces a novel yet simple pre-training strategy designed to guide a model away from learning domain-specific features in a pre-training dataset during pre-training, thereby improving the generalisation ability of the pre-trained model. To evaluate the strategy's effectiveness, deep learning models are pre-trained on ImageNet and subsequently fine-tuned on four semantic segmentation datasets with diverse scenes and modalities, including iSAID, MFNet, PST900 and Potsdam. Experimental results show that the proposed pre-training strategy led to state-of-the-art accuracies on all four datasets, namely 67.4% mIoU for iSAID, 56.9% mIoU for MFNet, 84.22% mIoU for PST900, 91.88% mF1 for Potsdam. This research lays the groundwork for developing a unified foundation model applicable to both computer vision and remote sensing applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。