arXiv:2511.01990cs.CV2025-11被引 4

GFMs在洪水淹没制图中表现优异,精度高且推理快,适合实际应用。

Assessing the value of Geo-Foundational Models for Flood Inundation Mapping: Benchmarking models for Sentinel-1, Sentinel-2, and Planetscope for end-users

  • 对比多种GFMs与传统模型,发现GFMs在不同卫星数据上均表现稳定
  • Clay模型在多传感器下综合性能最优,5张图训练即达0.64 mIoU
  • GFMs比U-Net更省算力和标注成本,适合资源有限的用户

Geo-Foundational Models (GFMs) 能快速可靠地从遥感影像中提取时空信息,提升洪水淹没制图效果。尽管潜力巨大,但其是否优于传统模型(如U-Net)仍不明确。为此,本文系统评估了Prithvi 2.0、Clay V1.5、DOFA及UViT(Prithvi变体)与TransNorm、U-Net、Attention U-Net在PlanetScope、Sentinel-1、Sentinel-2上的表现。结果显示,所有GFMs性能接近,跨传感器最佳与最差模型间仅相差2%-5%。Clay在PlanetScope(0.79 mIoU)和Sentinel-2(0.70 mIoU)上领先,而Prithvi在Sentinel-1上表现最佳(0.57 mIoU)。在五区域留一区域交叉验证中,Clay平均mIoU为0.72(0.04)/0.66(0.07)/0.51(0.08),优于Prithvi与DOFA。全19站点验证显示,Clay相较U-Net提升4%。视觉分析表明其更擅长保留细粒度特征。少样本实验中,仅需5张训练图,Clay即达0.64 mIoU,显著优于Prithvi(0.24)和DOFA(0.35)。Clay参数量仅26M,推理速度约为Prithvi(650M)的3倍、DOFA(410M)的2倍。结果表明,GFMs相比传统U-Net在洪水制图中提供小至中等精度提升,同时大幅降低计算与标注成本。

原文摘要 · Abstract (English)

Geo-Foundational Models (GFMs) enable fast and reliable extraction of spatiotemporal information from satellite imagery, improving flood inundation mapping by leveraging location and time embeddings. Despite their potential, it remains unclear whether GFMs outperform traditional models like U-Net. A systematic comparison across sensors and data availability scenarios is still lacking, which is an essential step to guide end-users in model selection. To address this, we evaluate three GFMs, Prithvi 2.0, Clay V1.5, DOFA, and UViT (a Prithvi variant), against TransNorm, U-Net, and Attention U-Net using PlanetScope, Sentinel-1, and Sentinel-2. We observe competitive performance among all GFMs, with only 2-5% variation between the best and worst models across sensors. Clay outperforms others on PlanetScope (0.79 mIoU) and Sentinel-2 (0.70), while Prithvi leads on Sentinel-1 (0.57). In leave-one-region-out cross-validation across five regions, Clay shows slightly better performance across all sensors (mIoU: 0.72(0.04), 0.66(0.07), 0.51(0.08)) compared to Prithvi (0.70(0.05), 0.64(0.09), 0.49(0.13)) and DOFA (0.67(0.07), 0.64(0.04), 0.49(0.09)) for PlanetScope, Sentinel-2, and Sentinel-1, respectively. Across all 19 sites, leave-one-region-out cross-validation reveals a 4% improvement by Clay compared to U-Net. Visual inspection highlights Clay's superior ability to retain fine details. Few-shot experiments show Clay achieves 0.64 mIoU on PlanetScope with just five training images, outperforming Prithvi (0.24) and DOFA (0.35). In terms of computational time, Clay is a better choice due to its smaller model size (26M parameters), making it ~3x faster than Prithvi (650M) and 2x faster than DOFA (410M). Contrary to previous findings, our results suggest GFMs offer small to moderate improvements in flood mapping accuracy at lower computational cost and labeling effort compared to traditional U-Net.

洪水制图遥感模型效率少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。