融合多源地理数据可显著提升卫星遥感模型的数据效率与泛化能力
Using Multiple Input Modalities Can Improve Data-Efficiency and O.O.D. Generalization for ML with Satellite Imagery
- 在卫星图像基础上加入地形、气象等多模态数据,提升模型性能
- 数据稀缺和跨区域场景下性能提升最明显,验证了多模态优势
- 硬编码融合策略优于学习型融合,对模型设计有新启示
全球范围内存在多种地理空间数据层,包括遥感栅格数据(如卫星影像、数字高程模型、预测土地覆盖图)以及环境传感器数据(如气温、风速)。然而,大多数训练于卫星影像的机器学习模型(SatML)主要依赖光学模态(如多光谱影像)。为探究在监督学习中结合其他输入模态的价值,我们通过在分类、回归和分割任务的数据集上附加地理数据层,构建了增强版的SatML基准任务。实验发现,将地理信息与光学影像融合可显著提升模型表现;该优势在标注数据有限及地理外样本场景下尤为突出,表明多模态输入对数据效率和跨区域泛化具有重要意义。令人意外的是,硬编码融合策略优于学习型融合方法,对后续研究具有重要启示。
原文摘要 · Abstract (English)
A large variety of geospatial data layers is available around the world ranging from remotely-sensed raster data like satellite imagery, digital elevation models, predicted land cover maps, and human-annotated data, to data derived from environmental sensors such as air temperature or wind speed data. A large majority of machine learning models trained on satellite imagery (SatML), however, are designed primarily for optical input modalities such as multi-spectral satellite imagery. To better understand the value of using other input modalities alongside optical imagery in supervised learning settings, we generate augmented versions of SatML benchmark tasks by appending additional geographic data layers to datasets spanning classification, regression, and segmentation. Using these augmented datasets, we find that fusing additional geographic inputs with optical imagery can significantly improve SatML model performance. Benefits are largest in settings where labeled data are limited and in geographic out-of-sample settings, suggesting that multi-modal inputs may be especially valuable for data-efficiency and out-of-sample performance of SatML models. Surprisingly, we find that hard-coded fusion strategies outperform learned variants, with interesting implications for future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。