利用多时相卫星图像提升变电站分割精度,效果优于传统数据增强
Improving satellite imagery segmentation using multiple Sentinel-2 revisits
- 在模型隐空间融合多时相图像特征,实现更优分割
- 基于SWIN Transformer的模型性能优于U-Net和ViT
- 结果在变电站分割与建筑密度估计任务中均验证有效
近年来,遥感数据分析受益于计算机视觉技术的引入,如使用在大规模多样化数据集上预训练的共享模型。然而,卫星影像具有传统计算机视觉未考虑的独特属性,例如同一地点的多次重访。本文探讨如何在微调预训练遥感模型的框架内最优利用这些重访数据。研究聚焦于气候减缓相关的实际问题——变电站分割,该问题可代表预训练模型的一般应用。通过在多种模型架构上测试不同多时相输入方案,发现将多个重访图像的特征在模型隐空间中融合,优于其他方法(包括作为数据增强)。此外,基于SWIN Transformer的架构表现优于U-Net和基于ViT的模型。我们在独立的建筑密度估计任务上验证了结果的普适性。
原文摘要 · Abstract (English)
In recent years, analysis of remote sensing data has benefited immensely from borrowing techniques from the broader field of computer vision, such as the use of shared models pre-trained on large and diverse datasets. However, satellite imagery has unique features that are not accounted for in traditional computer vision, such as the existence of multiple revisits of the same location. Here, we explore the best way to use revisits in the framework of fine-tuning pre-trained remote sensing models. We focus on an applied research question of relevance to climate change mitigation -- power substation segmentation -- that is representative of applied uses of pre-trained models more generally. Through extensive tests of different multi-temporal input schemes across diverse model architectures, we find that fusing representations from multiple revisits in the model latent space is superior to other methods of using revisits, including as a form of data augmentation. We also find that a SWIN Transformer-based architecture performs better than U-nets and ViT-based models. We verify the generality of our results on a separate building density estimation task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。