arXiv:2603.01576cs.CV2026-03被引 4

首个针对冰冻圈应用的模型评测基准,验证了大模型在冰雪场景中的泛化能力。

Cryo-Bench: Benchmarking Foundation Models for Cryosphere Applications

  • 构建覆盖冰川、海冰等多类冰冻圈任务的评测数据集,支持多传感器多区域评估。
  • 冻结编码器下UNet平均mIoU达66.38,少样本时大模型表现优于传统方法。
  • 推荐微调时优化学习率,可提升性能12.77%,适合快速部署或深入研究者使用。

地理基础模型(GFMs)已在多种地球观测任务中展现潜力,能在标签稀疏情况下生成可靠地图。然而,针对冰冻圈应用的评测仍受限于缺乏合适的数据集。为此,我们提出 extbf{Cryo-Bench},一个用于评估基础模型在关键冰冻圈成分上表现的基准,涵盖冰川碎屑覆盖区、冰湖、海冰及冰架断裂线,覆盖多种传感器与广泛地理区域。我们评估了14种GFMs及UNet、ViT基线模型。在编码器冻结条件下,UNet达到最高平均mIoU为66.38,次为TerraMind的64.02。在少样本设置(10%输入数据)下,DOFA和TerraMind分别取得59.53和56.62的mIoU,优于UNet的56.60。全量微调时模型表现不一致,但调整学习率显著提升性能:在GLID与CaFFe两个代表性数据集上,平均相对提升达12.77%。尽管预训练数据中冰冻圈信息极少,这些模型仍展现出显著领域适应能力,能有效完成各类任务。建议采用编码器微调结合超参数优化以获最优效果,而冻结编码器则适用于快速出结果。

原文摘要 · Abstract (English)

Geo-Foundation Models (GFMs) have been evaluated across diverse Earth observation task including multiple domains and have demonstrated strong potential of producing reliable maps even with sparse labels. However, benchmarking GFMs for Cryosphere applications has remained limited, primarily due to the lack of suitable evaluation datasets. To address this gap, we introduce \textbf{Cryo-Bench}, a benchmark compiled to evaluate GFM performance across key Cryospheric components. Cryo-Bench includes debris-covered glaciers, glacial lakes, sea ice, and calving fronts, spanning multiple sensors and broad geographic regions. We evaluate 14 GFMs alongside UNet and ViT baselines to assess their advantages, limitations, and optimal usage strategies. With a frozen encoder, UNet achieves the highest average mIoU of \textbf{66.38}, followed by TerraMind at \textbf{64.02} across five evluation dataset included in Cryo-Bench. In the few-shot setting (10\% input data), GFMs such as DOFA and TerraMind outperform UNet, achieving mIoU scores of \textbf{59.53}, \textbf{56.62}, and \textbf{56.60}, respectively, comapred to U-Net's 56.60. When fully finetuning GFMs, we observe inconsistent performance across datasets and models. However, tuning learning rate along with finetuning substantially improves GFM performance. For example, evaluation on two representative datasets (GLID and CaFFe) shows an average relative improvement of \textbf{12.77\%}. Despite having minimal Cryosphere representation in their pretraining data, GFMs exhibit notable domain adaptation capabilities and produce meaningful results across tasks. Based on our findings, We recommend encoder fine-tuning with hyperparameter optimization optimization to achieve the best possible performance, while using frozen encoders when users need quick results without extensive experimentation.(\href{https://github.com/Sk-2103/Cryo-Bench}{GitHub}).

冰冻圈基础模型评测基准少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。