arXiv:2506.09327cs.CV2025-06被引 5

利用多模态自监督学习,用少量标注数据提升高分辨率遥感图像理解能力。

MSSDF: Modality-Shared Self-supervised Distillation for High-Resolution Multi-modal Remote Sensing Image Learning

  • 设计跨模态掩码与自适应信息屏蔽策略,融合RGB、多光谱和数字地表模型。
  • 在15个数据集26项任务中多数超越现有方法,语义分割最高达78.30% mIoU。
  • 适合遥感图像分析、环境监测等需少样本预训练的科研与工程应用。

遥感图像解析在环境监测、城市规划和灾害评估中至关重要,但高质量标注数据获取成本高。本文提出一种多模态自监督学习框架,利用高分辨率RGB图像、多光谱数据和数字地表模型(DSM)进行预训练。通过设计信息感知的自适应掩码策略、跨模态掩码机制及多任务自监督目标,有效捕捉不同模态间的关联性与各模态内部的特征结构。我们在15个遥感数据集上进行了26项下游任务评估,涵盖场景分类、语义分割、变化检测、目标检测和深度估计。实验表明,该方法在多数任务上优于现有预训练方法。在Potsdam和Vaihingen语义分割任务中,仅使用50%训练集时,mIoU分别达到78.30%和76.50%;在US3D深度估计任务中,RMSE降至0.182;在SECOND数据集二值变化检测任务中,mIoU达47.51%,比第二名CS-MAE高出3个百分点。代码、模型权重和HR-Pairs数据集见https://github.com/CVEO/MSSDF。

原文摘要 · Abstract (English)

Remote sensing image interpretation plays a critical role in environmental monitoring, urban planning, and disaster assessment. However, acquiring high-quality labeled data is often costly and time-consuming. To address this challenge, we proposes a multi-modal self-supervised learning framework that leverages high-resolution RGB images, multi-spectral data, and digital surface models (DSM) for pre-training. By designing an information-aware adaptive masking strategy, cross-modal masking mechanism, and multi-task self-supervised objectives, the framework effectively captures both the correlations across different modalities and the unique feature structures within each modality. We evaluated the proposed method on multiple downstream tasks, covering typical remote sensing applications such as scene classification, semantic segmentation, change detection, object detection, and depth estimation. Experiments are conducted on 15 remote sensing datasets, encompassing 26 tasks. The results demonstrate that the proposed method outperforms existing pretraining approaches in most tasks. Specifically, on the Potsdam and Vaihingen semantic segmentation tasks, our method achieved mIoU scores of 78.30\% and 76.50\%, with only 50\% train-set. For the US3D depth estimation task, the RMSE error is reduced to 0.182, and for the binary change detection task in SECOND dataset, our method achieved mIoU scores of 47.51\%, surpassing the second CS-MAE by 3 percentage points. Our pretrain code, checkpoints, and HR-Pairs dataset can be found in https://github.com/CVEO/MSSDF.

遥感图像自监督学习多模态少样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。