arXiv:2512.02055cs.CVcs.AI2025-12被引 2

用多模态卫星数据微调通用模型,提升全球实时洪灾制图精度。

Leveraging AI multimodal geospatial foundation models for improved near-real-time flood mapping at a global scale

  • 融合Sentinel-1与Sentinel-2多源遥感数据,微调地球空间基础模型。
  • 基线未冻结配置在精度与效率间取得最佳平衡,召回率超基准模型。
  • 适用于灾害应急响应与气候适应决策,尤其适合缺乏标注数据地区。

洪水是破坏性最强的气象灾害之一,2024年作为有记录以来最热的一年,极端洪灾影响了五大洲多个社区。地球观测(EO)卫星提供关键且频繁的覆盖,用于洪涝范围制图,但实际精度严重依赖标注数据集和模型泛化能力。近期的地表空间基础模型(如ESA-IBM的TerraMind)通过大规模自监督预训练提升了泛化能力,但其在多样化全球洪灾场景下的表现仍不明确。本研究使用包含85个全球洪灾事件的多模态数据集FloodsNet(融合同址Sentinel-1 SAR与Sentinel-2光学影像),对TerraMind进行微调以实现洪涝范围分割。测试了四种配置(基础版与大模型;冻结与未冻结主干网络),并与TerraMind Sen1Floods11示例及在FloodsNet与Sen1Floods11上训练的U-Net对比。结果显示,基础版未冻结配置在准确率、精确率与召回率间取得最佳平衡,且计算成本显著低于大模型;大模型未冻结配置召回率最高。在相同整体准确率下,基于FloodsNet训练的模型召回率优于仅用Sen1Floods11训练的示例。尽管U-Net召回率更高,但其精确率和准确率略低。结果表明,融合多模态光学与雷达数据并微调基础模型,可有效提升近实时洪灾制图能力。本研究为首个针对全球尺度洪水分割的GFM评估,揭示其在气候适应与灾害韧性中的潜力与局限。

原文摘要 · Abstract (English)

Floods are among the most damaging weather-related hazards, and in 2024, the warmest year on record, extreme flood events affected communities across five continents. Earth observation (EO) satellites provide critical, frequent coverage for mapping inundation, yet operational accuracy depends heavily on labeled datasets and model generalization. Recent Geospatial Foundation Models (GFMs), such as ESA-IBM's TerraMind, offer improved generalizability through large-scale self-supervised pretraining, but their performance on diverse global flood events remains poorly understood. We fine-tune TerraMind for flood extent mapping using FloodsNet, a harmonized multimodal dataset containing co-located Sentinel-1 (Synthetic Aperture Radar, SAR data) and Sentinel-2 (optical) imagery for 85 flood events worldwide. We tested four configurations (base vs. large models; frozen vs. unfrozen backbones) and compared against the TerraMind Sen1Floods11 example and a U-Net trained on both FloodsNet and Sen1Floods11. The base-unfrozen configuration provided the best balance of accuracy, precision, and recall at substantially lower computational cost than the large model. The large unfrozen model achieved the highest recall. Models trained on FloodsNet outperformed the Sen1Floods11-trained example in recall with similar overall accuracy. U-Net achieved higher recall than all GFM configurations, though with slightly lower accuracy and precision. Our results demonstrate that integrating multimodal optical and SAR data and fine-tuning a GFM can enhance near-real-time flood mapping. This study provides one of the first global-scale evaluations of a GFM for flood segmentation, highlighting both its potential and current limitations for climate adaptation and disaster resilience.

洪灾制图多模态遥感基础模型实时监测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。