arXiv:2608.16614cs.CV2026-08

GeoFMs在真实场景中常过度自信,需用校准度评估其可靠性。

Beyond Accuracy: Assessing Calibration of Geospatial Foundation Models and Their Sensitivity to Distribution Shifts

论文配图:Beyond Accuracy: Assessing Calibration of Geospatial Foundation Models and Their Sensitivity to Distribution Shifts
图 1 · 摘自论文原文
  • 通过校准度分析发现,地理预训练模型在数据扰动下更易过度自信。
  • 16个编码器在多种数据集上均随干扰加剧而性能下降,排名变化显著。
  • 建议结合多条件、多指标评估,推动模型向实际部署靠拢。

地理空间基础模型(GeoFMs)通常仅依据标准基准下的平均准确率排名选择。我们指出该方法过于狭窄:在关键遥感任务部署中,需关注校准度——模型置信度与实际正确性的匹配程度。在16个冻结编码器、4个分类与5个分割数据集,以及两个正交扰动轴下,所有编码器在干扰增强时性能均下降,排名也随之变化。在4个分类基准中,经遥感预训练与图像网预训练的编码器在干净数据上的准确率与校准度无差异,且在分布偏移下稳定性相当。然而在偏移条件下,遥感预训练模型比图像网预训练模型更易陷入过度自信,各等级及各类扰动中皆如此。中心核对齐(CKA)分析表明,这源于表征刚性:遥感预训练嵌入在扰动下移动更少,但任务信息损失相同且仍保持高置信度。应用三种常见不确定性量化方法发现,温度缩放与深度集成无法缓解退化,而高斯过程探测器虽在严重云遮下将ECE减半,却使干净数据上的误差增至三倍。选择性预测实验显示,基于置信度的回避策略无法避免高自信错误预测。因此,我们主张评估应覆盖多重条件与指标,以更全面评估模型进展,缩小与真实部署的差距。

原文摘要 · Abstract (English)

Geospatial Foundation Models (GeoFMs) are most commonly ranked and selected by accuracy on standard benchmark conditions via averaged ranks. We show that this protocol is too narrow: the promised deployment in critical EO tasks requires further angles of analysis, mainly calibration, the agreement between a model's confidence and its correctness. Across 16 frozen encoders, four classification and five segmentation datasets, and two orthogonal stress axes, every encoder degrades as corruption intensifies, and the ranking changes as well. Across the four classification benchmarks, EO-pretrained and ImageNet-pretrained encoders are indistinguishable on clean accuracy and clean calibration, and EO pretraining provides no more stability under shift than ImageNet pretraining. Under shift the GeoFMs drift further into overconfidence than the ImageNet-pretrained encoders, at every grade and in every corruption family. A centered kernel alignment (CKA) analysis ties this to representational rigidity: EO-pretrained embeddings move less under corruption while losing just as much task information and remaining overconfident. We apply three commonly explored uncertainty quantification methods and find that temperature scaling and deep ensembles cannot counteract the degradation, while a Gaussian-process probe roughly halves ECE under severe cloud only by tripling it on clean data. In selective prediction experiments, we find that confidence-based abstention cannot defer around confidently wrong predictions, and advocate that benchmark rankings and evaluations should therefore operate across a multitude of conditions and metrics to more holistically evaluate model development progress and close the gap to real world deployment scenarios.

地理模型校准度遥感不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。