arXiv:2606.16302cs.CV2026-06

用Transformer比传统CNN更好分割卫星洪水影像,尤其在碎片化洪水区域表现更优。

Explainable Flood Segmentation on Sentinel-1 SAR1 Imagery Using CNN and Transformer Architectures

论文配图:Explainable Flood Segmentation on Sentinel-1 SAR1 Imagery Using CNN and Transformer Architectures
图 1 · 摘自论文原文
  • 对比了CNN与Vision Transformer在洪水分割中的表现,聚焦于区分洪水区与永久水体。
  • SegFormer-b2在ETCI数据集上洪水交并比显著优于U-Net,且结果更稳定。
  • 通过Grad-CAM和不确定性估计提升模型可解释性,适合灾害应急与遥感研究者使用。

快速准确的洪水预测对灾害响应与减灾规划至关重要。合成孔径雷达(SAR)卫星传感器可在全天候、无光照条件下工作,适用于洪水监测。然而,如何区分临时洪水区域与永久水体仍是挑战,尤其当洪水定义为被淹没的土地时。本研究系统比较了基于CNN的U-Net、U-Net++、DeepLabV3(ResNet-34主干)与三种SegFormer变体(b0、b1、b2)在Sentinel-1 SAR影像上的多类洪水分割性能,在ETCI NASA与Sen1Floods11两个基准数据集上采用场景级划分进行评估,以检验空间泛化能力。结果表明,SegFormer-b2在ETCI数据集上整体显著优于U-Net(Wilcoxon符号秩检验下7个测试场景均更高),而在经过Sen1Floods11微调后,优势缩小至场景变异范围内,主要集中在空间碎片化的洪水事件中。研究结合定性与定量可解释性方法,可视化模型决策过程并评估预测可靠性:SegFormer-b2生成的Grad-CAM激活更集中于洪水相关特征,而U-Net在洪水边界处提供更丰富的不确定性估计。

原文摘要 · Abstract (English)

Rapid and accurate flood prediction is essential for disaster response and mitigation planning. Synthetic Aperture Radar (SAR) sensors in satellites are well-suited for this purpose because they operate independently of weather and daylight conditions. Although SAR-based data enable all-weather flood monitoring, distinguishing flooded land from permanent water remains a significant challenge, particularly when flooding is defined strictly as inundated land. This study provides a comprehensive comparison of convolutional neural network (CNN) and vision transformer architectures for multi-class flood segmentation using Sentinel-1 SAR imagery, specifically trained to separate flooded land from permanent water bodies and land. Three state-of-the-art (SOTA)CNN-based models, U-Net, U-Net++, and DeepLabV3 with ResNet-34 backbone, and three SegFormer variants (b0,b1,b2) were evaluated in two benchmark datasets, the ETCI NASA dataset and SenFloods11, using scene-based data splits to ensure a realistic assessment of spatial generalization. The results demonstrate that SegFormer-b2 significantly outperforms the U-Net baseline on the ETCI dataset (higher flood IoU across all 7 test scenes in the Wilcoxon signed-rank test), while after fine-tuning on Sen1Floods11, the advantage narrows to within the range of scene variability and is concentrated in spatially fragmented flood events. The study includes both qualitative and quantitative explainability techniques to visually comprehend model decisions and systematically assess prediction reliability. Qualitative analysis reveals that SegFormer-b2 produces more spatially coherent Grad-CAM activations focused on flood-relevant features, while U-Net generates more informative uncertainty estimates along flood boundaries.

洪水分割遥感图像Transformer可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。