用位置预测提升多模态卫星图像分割效果,解决标注数据少的难题。
Position Prediction Self-Supervised Learning for Multimodal Satellite Imagery Semantic Segmentation
- 通过相对位置预测替代重建,增强空间定位能力
- 在Sen1Floods11数据集上超越现有自监督方法
- 适合需要少标注数据的遥感图像分割任务
卫星影像语义分割对地球观测至关重要,但受限于标注数据稀缺。现有自监督方法如掩码自编码器(MAE)侧重重建而非定位,而定位是分割的核心。本文将位置感知的LOCA方法适配至多模态卫星影像,通过扩展SatMAE的通道分组至多模态数据,并引入同组注意力掩码促进跨模态交互。采用相对块位置预测,推动模型学习空间推理能力。在Sen1Floods11洪水制图数据集上,该方法显著优于基于重建的自监督方法,证明位置预测任务经合理适配后,能生成更利于卫星影像分割的表征。
原文摘要 · Abstract (English)
Semantic segmentation of satellite imagery is crucial for Earth observation applications, but remains constrained by limited labelled training data. While self-supervised pretraining methods like Masked Autoencoders (MAE) have shown promise, they focus on reconstruction rather than localisation-a fundamental aspect of segmentation tasks. We propose adapting LOCA (Location-aware), a position prediction self-supervised learning method, for multimodal satellite imagery semantic segmentation. Our approach addresses the unique challenges of satellite data by extending SatMAE's channel grouping from multispectral to multimodal data, enabling effective handling of multiple modalities, and introducing same-group attention masking to encourage cross-modal interaction during pretraining. The method uses relative patch position prediction, encouraging spatial reasoning for localisation rather than reconstruction. We evaluate our approach on the Sen1Floods11 flood mapping dataset, where it significantly outperforms existing reconstruction-based self-supervised learning methods for satellite imagery. Our results demonstrate that position prediction tasks, when properly adapted for multimodal satellite imagery, learn representations more effective for satellite image semantic segmentation than reconstruction-based approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。