用自监督预训练提升ALOS-2 SAR图像语义分割效果
A Tutorial on ALOS2 SAR Utilization: Dataset Preparation, Self-Supervised Pretraining, and Semantic Segmentation
- 提出SAR-W-SimMIM,加权处理斑点噪声和极端强度值
- 在日本区域数据集上预训练,分割性能显著优于随机初始化
- 为构建全国尺度遥感基础模型提供数据与方法指南
掩码自编码器(MAE)等方法在卫星影像中表现良好,但应用于合成孔径雷达(SAR)仍受限于语义标注难和高噪声问题。基于先前的SAR-W-MixMAE工作,本文引入SAR-W-SimMIM,一种针对ALOS-2单通道(HH极化)SAR影像的加权版SimMIM方法,旨在降低斑点噪声和极端强度值对自监督预训练的影响。在与SAR-W-MixMAE及随机初始化对比中,该方法在语义分割任务上取得显著提升。此外,由于地表覆盖分布不均(如水体、森林或沙漠占主导),区域特异性模型的预训练与微调面临偏差问题。为此,本文构建了面向日本地区的ALOS-2单通道SAR数据集,作为全国尺度基础模型的初步尝试。使用视觉变换器架构的自编码器在此数据集上预训练,并通过任务特定解码器微调用于语义分割。初步结果表明,相比从头训练,性能有明显提升。本工作系统介绍了如何处理与准备ALOS-2观测数据,以支持自监督预训练和下游任务(如语义分割)的微调。
原文摘要 · Abstract (English)
Masked auto-encoders (MAE) and related approaches have shown promise for satellite imagery, but their application to synthetic aperture radar (SAR) remains limited due to challenges in semantic labeling and high noise levels. Building on our prior work with SAR-W-MixMAE, which adds SAR-specific intensity-weighted loss to standard MixMAE for pretraining, we also introduce SAR-W-SimMIM; a weighted variant of SimMIM applied to ALOS-2 single-channel SAR imagery. This method aims to reduce the impact of speckle and extreme intensity values during self-supervised pretraining. We evaluate its effect on semantic segmentation compared to our previous trial with SAR-W-MixMAE and random initialization, observing notable improvements. In addition, pretraining and fine-tuning models on satellite imagery pose unique challenges, particularly when developing region-specific models. Imbalanced land cover distributions such as dominant water, forest, or desert areas can introduce bias, affecting both pretraining and downstream tasks like land cover segmentation. To address this, we constructed a SAR dataset using ALOS-2 single-channel (HH polarization) imagery focused on the Japan region, marking the initial phase toward a national-scale foundation model. This dataset was used to pretrain a vision transformer-based autoencoder, with the resulting encoder fine-tuned for semantic segmentation using a task-specific decoder. Initial results demonstrate significant performance improvements compared to training from scratch with random initialization. In summary, this work provides a guide to process and prepare ALOS2 observations to create dataset so that it can be taken advantage of self-supervised pretraining of models and finetuning downstream tasks such as semantic segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。