构建全球多元地理空间模型评估基准,揭示现有模型在真实场景下的局限性。
PANGAEA: A Global and Inclusive Benchmark for Geospatial Foundation Models
- 提出覆盖多分辨率、多传感器、多时相的标准化评估协议
- 实测主流模型在少样本下表现不及监督模型,暴露泛化瓶颈
- 支持持续扩展,助力地理空间大模型可信评估,适合遥感与地球科学研究者
地理空间基础模型(GFMs)已成为从地球观测数据中提取表征的强大工具,但其评估仍存在不一致且范围狭窄的问题。现有工作常在次优的下游数据集和任务上评估,这些任务往往过于简单或过于狭隘,难以真实反映模型的实际应用能力。此外,当前评估协议缺乏多样性,未能涵盖图像分辨率、传感器类型和时间维度的多样性,进一步阻碍了对模型性能的全面评估。尤其多数基准地理上偏向北美和欧洲,质疑了模型的全球适用性。为此,我们提出PANGAEA,一个涵盖多样数据集、任务、分辨率、传感器模态和时间跨度的标准评估协议,建立稳健且广泛适用的基准。我们在该基准上评估了最流行的开源GFMs,并分析其在多个领域中的表现。特别地,将这些模型与监督基线(如UNet和普通ViT)进行比较,评估其在有限标注数据下的有效性。结果表明,在不同场景下GFMs并不始终优于监督模型。PANGAEA设计为高度可扩展,未来可无缝集成新数据集、模型和任务。通过发布评估代码和基准,我们希望促进研究者复现实验并推动更严谨的大地理空间模型评估范式。代码见https://github.com/VMarsocci/pangaea-bench。
原文摘要 · Abstract (English)
Geospatial Foundation Models (GFMs) have emerged as powerful tools for extracting representations from Earth observation data, but their evaluation remains inconsistent and narrow. Existing works often evaluate on suboptimal downstream datasets and tasks, that are often too easy or too narrow, limiting the usefulness of the evaluations to assess the real-world applicability of GFMs. Additionally, there is a distinct lack of diversity in current evaluation protocols, which fail to account for the multiplicity of image resolutions, sensor types, and temporalities, which further complicates the assessment of GFM performance. In particular, most existing benchmarks are geographically biased towards North America and Europe, questioning the global applicability of GFMs. To overcome these challenges, we introduce PANGAEA, a standardized evaluation protocol that covers a diverse set of datasets, tasks, resolutions, sensor modalities, and temporalities. It establishes a robust and widely applicable benchmark for GFMs. We evaluate the most popular GFMs openly available on this benchmark and analyze their performance across several domains. In particular, we compare these models to supervised baselines (e.g. UNet and vanilla ViT), and assess their effectiveness when faced with limited labeled data. Our findings highlight the limitations of GFMs, under different scenarios, showing that they do not consistently outperform supervised models. PANGAEA is designed to be highly extensible, allowing for the seamless inclusion of new datasets, models, and tasks in future research. By releasing the evaluation code and benchmark, we aim to enable other researchers to replicate our experiments and build upon our work, fostering a more principled evaluation protocol for large pre-trained geospatial models. The code is available at https://github.com/VMarsocci/pangaea-bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。