arXiv:2603.02522cs.CV2026-03被引 1

利用邻接遥感图像的空间关联性提升自监督表征学习效果

NeighborMAE: Exploiting Spatial Dependencies between Neighboring Earth Observation Images in Masked Autoencoders Pretraining

  • 通过联合重建相邻遥感图像来捕捉空间依赖关系
  • 在多个数据集和任务上显著优于现有基线方法
  • 适合遥感图像分析、地理信息建模等应用

掩码图像建模是利用大规模未标注遥感图像进行自监督表征学习的重要范式。尽管多模态和多时相遥感数据已广泛融入该框架,但相邻区域图像间的空间依赖关系仍被忽视。由于地表具有连续性,相邻图像高度相关,蕴含丰富的上下文信息。为此,我们提出NeighborMAE,通过联合重建邻近遥感图像来学习空间依赖关系。为保持重建挑战性,采用启发式策略动态调整掩码比例与像素级损失权重。在多个预训练数据集和下游任务上的实验结果表明,NeighborMAE显著优于现有基线,证明了邻接图像在遥感掩码建模中的价值及所提设计的有效性。

原文摘要 · Abstract (English)

Masked Image Modeling has been one of the most popular self-supervised learning paradigms to learn representations from large-scale, unlabeled Earth Observation images. While incorporating multi-modal and multi-temporal Earth Observation data into Masked Image Modeling has been widely explored, the spatial dependencies between images captured from neighboring areas remains largely overlooked. Since the Earth's surface is continuous, neighboring images are highly related and offer rich contextual information for self-supervised learning. To close this gap, we propose NeighborMAE, which learns spatial dependencies by joint reconstruction of neighboring Earth Observation images. To ensure that the reconstruction remains challenging, we leverage a heuristic strategy to dynamically adjust the mask ratio and the pixel-level loss weight. Experimental results across various pretraining datasets and downstream tasks show that NeighborMAE significantly outperforms existing baselines, underscoring the value of neighboring images in Masked Image Modeling for Earth Observation and the efficacy of our designs.

遥感图像自监督学习空间依赖掩码建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。