arXiv:2507.16849cs.CVcs.AI2025-07被引 1

用Transformer模型提升灾后区域分割精度,支持灾情快速评估

Post-Disaster Affected Area Segmentation with a Vision Transformer (ViT)-based EVAP Model using Sentinel-2 and Formosat-5 Imagery

  • 基于少量人工标注,通过主成分分析扩展标签
  • 融合哨兵-2与福卫五号多光谱数据训练模型
  • 适合缺乏真实标注的灾害应急测绘场景

本文提出一种基于视觉变换器(ViT)的深度学习框架,用于改进遥感影像中灾后受影响区域的分割,以支持台湾太空中心(TASA)开发的紧急增值产品(EVAP)。该方法从少量手动标注区域出发,利用主成分分析(PCA)进行特征空间分析,并构建置信度指数(CI)来扩展标签,生成弱监督训练集。随后,使用哨兵-2和福卫五号的多波段数据,训练基于ViT的编码器-解码器模型。架构支持多种解码器变体和多阶段损失策略,在监督有限的情况下提升性能。评估时,将模型输出与高分辨率EVAP结果对比,检验空间连贯性与分割一致性。在2022年鄱阳湖干旱与2023年罗德岛野火的案例研究中,本框架显著提升了分割结果的平滑性与可靠性,为缺乏真实标注数据的灾害制图提供可扩展方案。

原文摘要 · Abstract (English)

We propose a vision transformer (ViT)-based deep learning framework to refine disaster-affected area segmentation from remote sensing imagery, aiming to support and enhance the Emergent Value Added Product (EVAP) developed by the Taiwan Space Agency (TASA). The process starts with a small set of manually annotated regions. We then apply principal component analysis (PCA)-based feature space analysis and construct a confidence index (CI) to expand these labels, producing a weakly supervised training set. These expanded labels are then used to train ViT-based encoder-decoder models with multi-band inputs from Sentinel-2 and Formosat-5 imagery. Our architecture supports multiple decoder variants and multi-stage loss strategies to improve performance under limited supervision. During the evaluation, model predictions are compared with higher-resolution EVAP output to assess spatial coherence and segmentation consistency. Case studies on the 2022 Poyang Lake drought and the 2023 Rhodes wildfire demonstrate that our framework improves the smoothness and reliability of segmentation results, offering a scalable approach for disaster mapping when accurate ground truth is unavailable.

灾害监测遥感分割视觉Transformer弱监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。