arXiv:2603.22768cs.CV2026-03被引 1

用AI提升卫星图损伤识别,让灾后救援更精准。

From Pixels to Semantics: A Multi-Stage AI Framework for Structural Damage Detection in Satellite Imagery

  • 先超分再检测,结合视觉语言模型分四级判断损毁
  • 在xBD数据集上,4096×4096图像提升细节可见性
  • 无需真实标注,用CLIPScore和多模型投票保障可靠性

灾后快速准确评估建筑损伤对应急响应至关重要。然而,遥感影像常因空间分辨率低、上下文模糊和语义解析难,影响传统检测流程的可靠性。本文提出一种融合AI超分辨率、深度学习目标检测与视觉语言模型(VLM)的多阶段框架。首先,采用视频修复变换器(VRT)将灾前灾后卫星图像从1024×1024提升至4096×4096,增强结构细节可视性;其次,基于YOLOv11的检测器定位建筑,再对裁剪区域使用VLM进行四等级语义损伤分析;为在缺乏真实标注时确保评估稳健性,采用无参考的CLIPScore进行语义对齐,并引入多模型VLM-as-a-Jury策略降低单模型偏差。在xBD数据集的摩尔龙卷风和马修飓风子集上的实验表明,该框架显著提升了建筑损伤的语义理解能力,同时为一线救援人员提供恢复建议。

原文摘要 · Abstract (English)

Rapid and accurate structural damage assessment following natural disasters is critical for effective emergency response and recovery. However, remote sensing imagery often suffers from low spatial resolution, contextual ambiguity, and limited semantic interpretability, reducing the reliability of traditional detection pipelines. In this work, we propose a novel hybrid framework that integrates AI-based super-resolution, deep learning object detection, and Vision-Language Models (VLMs) for comprehensive post-disaster building damage assessment. First, we enhance pre- and post-disaster satellite imagery using a Video Restoration Transformer (VRT) to upscale images from 1024x1024 to 4096x4096 resolution, improving structural detail visibility. Next, a YOLOv11-based detector localizes buildings in pre-disaster imagery, and cropped building regions are analyzed using VLMs to semantically assess structural damage across four severity levels. To ensure robust evaluation in the absence of ground-truth captions, we employ CLIPScore for reference-free semantic alignment and introduce a multi-model VLM-as-a-Jury strategy to reduce individual model bias in safety-critical decision making. Experiments on subsets of the xBD dataset, including the Moore Tornado and Hurricane Matthew events, demonstrate that the proposed framework enhances the semantic interpretation of damaged buildings. In addition, our framework provides helpful recommendations to first responders for recovery based on damage analysis.

灾害评估卫星影像视觉语言模型损伤检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。