用一个模型搞定卫星洪水测绘,不管有啥传感器都能用。
Sensor-Adaptive Flood Mapping with Pre-trained Multi-Modal Transformers across SAR and Multispectral Modalities
- 轻量级多模态模型同时处理雷达与光学数据
- 融合模式下F1达0.896,单源也保持高效
- 适合灾情紧急时传感器不全的应急场景
洪水频发带来巨大人财物损失,亟需快速精准的淹没范围制图。尽管遥感技术提升了监测能力,但单一传感器受天气和重访周期限制,多源融合又需大量算力和标注数据。本文提出一种新型传感器自适应方法,通过微调Presto——一个仅约0.4M参数的轻量级多模态预训练变换器,在像素级处理合成孔径雷达(SAR)与多光谱(MS)数据。该模型可仅用SAR、仅用MS,或两者结合输入,统一实现洪水制图,满足灾害发生后第一时间使用可用数据的需求。在Sen1Floods11数据集上,相比大型基准模型Prithvi-100M(约100M参数),本方法在最优融合场景下取得F1分数0.896、mIoU 0.886,显著领先;在仅含多光谱数据时仍保持F1 0.893,仅用雷达数据时也能达到F1 0.718,证明了多模态预训练在实际应急场景中的鲁棒性。该参数高效、传感器灵活的方法为应对真实灾害中传感器受限的问题提供了可行方案。
原文摘要 · Abstract (English)
Floods are increasingly frequent natural disasters causing extensive human and economic damage, highlighting the critical need for rapid and accurate flood inundation mapping. While remote sensing technologies have advanced flood monitoring capabilities, operational challenges persist: single-sensor approaches face weather-dependent data availability and limited revisit periods, while multi-sensor fusion methods require substantial computational resources and large-scale labeled datasets. To address these limitations, this study introduces a novel sensor-flexible flood detection methodology by fine-tuning Presto, a lightweight ($\sim$0.4M parameters) multi-modal pre-trained transformer that processes both Synthetic Aperture Radar (SAR) and multispectral (MS) data at the pixel level. Our approach uniquely enables flood mapping using SAR-only, MS-only, or combined SAR+MS inputs through a single model architecture, addressing the critical operational need for rapid response with whatever sensor data becomes available first during disasters. We evaluated our method on the Sen1Floods11 dataset against the large-scale Prithvi-100M baseline ($\sim$100M parameters) across three realistic data availability scenarios. The proposed model achieved superior performance with an F1 score of 0.896 and mIoU of 0.886 in the optimal sensor-fusion scenario, outperforming the established baseline. Crucially, the model demonstrated robustness by maintaining effective performance in MS-only scenarios (F1: 0.893) and functional capabilities in challenging SAR-only conditions (F1: 0.718), confirming the advantage of multi-modal pre-training for operational flood mapping. Our parameter-efficient, sensor-flexible approach offers an accessible and robust solution for real-world disaster scenarios requiring immediate flood extent assessment regardless of sensor availability constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。