构建地理空间AI评估新框架,区分模型能力短板。
GEO-Bench-2: From Performance to Capability, Rethinking Evaluation in Geospatial AI
- 按分辨率、波段等特征分组评估,识别模型专长
- 多光谱任务中专用模型优于通用图像预训练模型
- 适合关注地理空间模型选型与改进的研究者
地理空间基础模型(GeoFMs)正重塑地球观测领域,但评估缺乏统一标准。GEO-Bench-2 提出覆盖分类、分割、回归、目标检测和实例分割的综合性评估框架,涵盖19个开源数据集。通过定义‘能力’分组,对具有相似特征(如分辨率、波段、时序性)的数据集进行模型排名,帮助用户识别各模型在特定能力上的表现,并定位未来研究需改进的方向。该框架提供可遵循但灵活的评估协议,既保证基准测试一致性,又支持模型适配策略研究,是推动GeoFMs在下游任务中发展的关键挑战。实验表明,无单一模型在所有任务中占优,架构设计与预训练选择具有任务特异性:在高分辨率任务中,自然图像预训练模型(ConvNext ImageNet, DINO V3)表现更优;而在农业、灾害响应等多光谱任务中,专用模型(TerraMind, Prithvi, Clay)显著领先。结果表明最优模型选择取决于任务需求、数据模态和约束条件。这说明实现跨任务通用的GeoFM仍是一个开放问题。GEO-Bench-2 支持面向具体应用场景的可复现评估。代码、数据与排行榜已公开发布。
原文摘要 · Abstract (English)
Geospatial Foundation Models (GeoFMs) are transforming Earth Observation (EO), but evaluation lacks standardized protocols. GEO-Bench-2 addresses this with a comprehensive framework spanning classification, segmentation, regression, object detection, and instance segmentation across 19 permissively-licensed datasets. We introduce ''capability'' groups to rank models on datasets that share common characteristics (e.g., resolution, bands, temporality). This enables users to identify which models excel in each capability and determine which areas need improvement in future work. To support both fair comparison and methodological innovation, we define a prescriptive yet flexible evaluation protocol. This not only ensures consistency in benchmarking but also facilitates research into model adaptation strategies, a key and open challenge in advancing GeoFMs for downstream tasks. Our experiments show that no single model dominates across all tasks, confirming the specificity of the choices made during architecture design and pretraining. While models pretrained on natural images (ConvNext ImageNet, DINO V3) excel on high-resolution tasks, EO-specific models (TerraMind, Prithvi, and Clay) outperform them on multispectral applications such as agriculture and disaster response. These findings demonstrate that optimal model choice depends on task requirements, data modalities, and constraints. This shows that the goal of a single GeoFM model that performs well across all tasks remains open for future research. GEO-Bench-2 enables informed, reproducible GeoFM evaluation tailored to specific use cases. Code, data, and leaderboard for GEO-Bench-2 are publicly released under a permissive license.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。