arXiv:2602.08117cs.CV2026-02被引 2

用卫星图+小模型检测灾后建筑损伤,效果优于传统方法。

Building Damage Detection using Satellite Images and Patch-Based Transformer Methods

论文配图:Building Damage Detection using Satellite Images and Patch-Based Transformer Methods
图 1 · 摘自论文原文
  • 通过裁片预处理提取结构特征,降低背景干扰。
  • 在噪声多、类别不平衡数据上实现接近最优的F1分数。
  • 适合灾害应急响应与遥感智能分析研究者参考。

快速建筑损伤评估对灾后响应至关重要。基于卫星影像的损伤分类模型可规模化获取态势感知信息。然而,卫星数据中的标签噪声和严重的类别不平衡带来了重大挑战。xBD数据集为跨地理区域的建筑级损伤提供了标准化基准。本研究评估了视觉变换器(ViT)在xBD数据集上的表现,特别关注其在噪声和不平衡数据下区分不同类型结构损伤的能力。我们具体评估了DINOv2-small和DeiT模型在多类损伤分类中的表现。提出一种针对裁片的预处理流程,以分离结构特征并最小化训练中的背景噪声。采用冻结头部微调策略,保持计算成本可控。模型性能通过准确率、精确率、召回率及宏平均F1分数进行评估。结果表明,结合新型训练方法的小型ViT架构,在宏平均F1上达到与先前卷积神经网络基线相当的水平,适用于灾难分类任务。

原文摘要 · Abstract (English)

Rapid building damage assessment is critical for post-disaster response. Damage classification models built on satellite imagery provide a scalable means of obtaining situational awareness. However, label noise and severe class imbalance in satellite data create major challenges. The xBD dataset offers a standardized benchmark for building-level damage across diverse geographic regions. In this study, we evaluate Vision Transformer (ViT) model performance on the xBD dataset, specifically investigating how these models distinguish between types of structural damage when training on noisy, imbalanced data. In this study, we specifically evaluate DINOv2-small and DeiT for multi-class damage classification. We propose a targeted patch-based pre-processing pipeline to isolate structural features and minimize background noise in training. We adopt a frozen-head fine-tuning strategy to keep computational requirements manageable. Model performance is evaluated through accuracy, precision, recall, and macro-averaged F1 scores. We show that small ViT architectures with our novel training method achieves competitive macro-averaged F1 relative to prior CNN baselines for disaster classification.

建筑损伤卫星图像Transformer灾害响应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。