arXiv:2508.17130cs.CV2025-08ICML

用AI提升灾后航拍画质并自动判损,准确率达84.5%

Structural Damage Detection Using AI Super Resolution and Visual Language Model

  • 融合超分辨率与视觉语言模型,从低清影像中识别损伤
  • 在土耳其地震与俄克拉荷马龙卷风数据上达84.5%分类准确率
  • 非专业用户也可操作,适合应急响应与资源有限地区

自然灾害突发性强、影响范围广,传统人工评估耗时费力且存在安全风险,难以满足快速响应需求,尤其在资源匮乏地区。本研究提出一种低成本新框架,结合无人机航拍影像、先进的视频超分辨率模型Video Restoration Transformer(VRT)和270亿参数的视觉语言模型Gemma3:27b。该系统可增强低分辨率灾后影像,自动识别结构损伤,并将建筑分为四类损坏等级(无/轻微损伤至完全破坏),同时输出对应风险等级。方法在2023年土耳其地震(数据来自《卫报》)及2013年摩尔龙卷风卫星数据(xBD数据集)上验证,整体分类准确率达84.5%,展现出高精度。系统具备良好可及性,使非技术人员也能完成初步分析,显著提升灾害管理响应速度与效率。

原文摘要 · Abstract (English)

Natural disasters pose significant challenges to timely and accurate damage assessment due to their sudden onset and the extensive areas they affect. Traditional assessment methods are often labor-intensive, costly, and hazardous to personnel, making them impractical for rapid response, especially in resource-limited settings. This study proposes a novel, cost-effective framework that leverages aerial drone footage, an advanced AI-based video super-resolution model, Video Restoration Transformer (VRT), and Gemma3:27b, a 27 billion parameter Visual Language Model (VLM). This integrated system is designed to improve low-resolution disaster footage, identify structural damage, and classify buildings into four damage categories, ranging from no/slight damage to total destruction, along with associated risk levels. The methodology was validated using pre- and post-event drone imagery from the 2023 Turkey earthquakes (courtesy of The Guardian) and satellite data from the 2013 Moore Tornado (xBD dataset). The framework achieved a classification accuracy of 84.5%, demonstrating its ability to provide highly accurate results. Furthermore, the system's accessibility allows non-technical users to perform preliminary analyses, thereby improving the responsiveness and efficiency of disaster management efforts.

灾后评估视觉语言模型图像超分无人机应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。