arXiv:2505.19447eess.IVcs.CV2025-05中稿 · publication in Geo…被引 2

通过精准对齐的图像对,实现遥感图像自监督预训练新方法。

A Contrastive Learning Foundation Model Based on Perfectly Aligned Sample Pairs for Remote Sensing Images

  • 利用空间不重叠掩码生成语义对齐样本对,提升特征一致性。
  • 在500万张未标注遥感图像上训练,性能媲美先进模型且显存更低。
  • 适合处理无标注、有噪声的遥感数据,推动实际应用落地。

自监督学习(SSL)可在无需昂贵标注数据的情况下预训练基础模型。对比学习(CL)在噪声干扰下更擅长获取精确的语义表示,但因域差异显著,在遥感(RS)图像任务中仍需特定适配。为此,本文提出一种新方法PerA,通过语义完全对齐的样本对生成通用遥感特征。PerA对增强图像施加空间不重叠掩码以采样视图,而非随机裁剪;通过确保教师-学生间的一致性并预测可学习掩码标记,提供高质量特征。相比以往对比方法,本方法具有更高内存效率,支持更大批次训练。此外,所提方法对未经清理的遥感数据表现出卓越适应性,并减轻潜在语义不一致的影响。我们还构建了一个包含约500万张遥感图像的无标签预训练数据集。在多个下游任务数据集上实验表明,仅用较小模型规模即达到与先前最优方法相当的性能,验证了该方法的有效性。希望本工作能助力遥感图像解析的实际应用。

原文摘要 · Abstract (English)

Self-Supervised Learning (SSL) enables us to pre-train foundation models without costly labeled data. Among SSL methods, Contrastive Learning (CL) methods are better at obtaining accurate semantic representations in noise interference. However, due to the significant domain gap, while CL methods have achieved great success in many computer vision tasks, they still require specific adaptation for Remote Sensing (RS) images. To this end, we present a novel self-supervised method called PerA, which produces all-purpose RS features through semantically Perfectly Aligned sample pairs. Specifically, PerA obtains features from sampled views by applying spatially disjoint masks to augmented images rather than random cropping. Our framework provides high-quality features by ensuring consistency between teacher and student and predicting learnable mask tokens. Compared to previous contrastive methods, our method demonstrates higher memory efficiency and can be trained with larger batches due to its sparse inputs. Additionally, the proposed method demonstrates remarkable adaptability to uncurated RS data and reduce the impact of the potential semantic inconsistency. We also collect an unlabeled pre-training dataset, which contains about 5 million RS images. We conducted experiments on multiple downstream task datasets and achieved performance comparable to previous state-of-the-art methods with a limited model scale, demonstrating the effectiveness of our approach. We hope this work will contribute to practical remote sensing interpretation works.

遥感图像自监督学习对比学习基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。