arXiv:2409.08744cs.CVcs.LG2024-09

评估遥感大模型在不同区域的泛化能力与不确定性,指导实际应用决策。

Uncertainty and Generalizability in Foundation Models for Earth Observation

  • 通过大规模消融实验对比8个大模型在11个区域的表现
  • 部分任务相关系数超0.9,但性能和不确定性差异显著
  • 建议用全局标签+简单探针方法优化下游任务设计

我们关注如何在有限标注预算下,针对特定兴趣区域(AOI)设计下游任务(如植被覆盖率估计)。利用现有基础模型(FM),需决定是基于标签丰富的其他区域训练并泛化至目标区域,还是在目标区域内部划分样本进行训练与验证。无论哪种方式,选择何种模型、如何采样标注数据等决策都会影响结果性能与不确定性。本文在哨兵1/2数据上,使用8个现有基础模型,以ESA世界覆盖产品类别为下游任务,在11个不同AOI上开展大规模消融研究。通过重复采样与训练,共构建约50万组简单线性回归模型。结果表明,跨区域空间泛化存在明显局限,但在特定任务中仍可实现预测值与真实值相关系数超过0.9;然而,性能与不确定性在不同区域、任务和模型间波动极大。我们认为这是实践中关键问题:每个基础模型与下游任务背后涉及大量设计选择(输入模态、采样策略、架构、预训练方式等),而下游任务设计者通常只能控制其中少数几项。本工作倡导在发布新基础模型时,以及在设计下游任务时,采用本文所述的大规模消融分析方法(基于参考全球标签与简单探针)来做出更明智的决策。

原文摘要 · Abstract (English)

We take the perspective in which we want to design a downstream task (such as estimating vegetation coverage) on a certain area of interest (AOI) with a limited labeling budget. By leveraging an existing Foundation Model (FM) we must decide whether we train a downstream model on a different but label-rich AOI hoping it generalizes to our AOI, or we split labels in our AOI for training and validating. In either case, we face choices concerning what FM to use, how to sample our AOI for labeling, etc. which affect both the performance and uncertainty of the results. In this work, we perform a large ablative study using eight existing FMs on either Sentinel 1 or Sentinel 2 as input data, and the classes from the ESA World Cover product as downstream tasks across eleven AOIs. We do repeated sampling and training, resulting in an ablation of some 500K simple linear regression models. Our results show both the limits of spatial generalizability across AOIs and the power of FMs where we are able to get over 0.9 correlation coefficient between predictions and targets on different chip level predictive tasks. And still, performance and uncertainty vary greatly across AOIs, tasks and FMs. We believe this is a key issue in practice, because there are many design decisions behind each FM and downstream task (input modalities, sampling, architectures, pretraining, etc.) and usually a downstream task designer is aware of and can decide upon a few of them. Through this work, we advocate for the usage of the methodology herein described (large ablations on reference global labels and simple probes), both when publishing new FMs, and to make informed decisions when designing downstream tasks to use them.

遥感基础模型不确定性泛化性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。