用多模态注意力模型自动识别灾后建筑损毁等级,提升应急响应效率。
Multi-Modal Attention for Automated Disaster Damage Assessment Using Remote Sensing Imagery and Deep Learning

- 通过跨模态注意力融合灾前灾后遥感影像特征,精准捕捉结构变化。
- 在大规模数据集上实现94.90%的整体分类准确率,对损坏类别区分力强。
- 适合灾害应急、城市规划与遥感分析人员快速部署使用。
及时准确的灾害损毁评估对有效应急响应、资源调配和恢复至关重要。传统方法常依赖人工巡查或稀疏数据,速度慢且易出错。本文提出一种基于遥感影像与深度学习的自动化建筑损毁分类框架。利用灾前灾后卫星影像,将建筑物分为无损、轻损、重损和倒塌四类。核心创新在于多模态注意力机制,可显式检测并评估结构变化。采用轻量级ConvNeXT-Tiny主干网络,在不损失性能的前提下保障高效处理。主要贡献包括:(1)用于多模态数据融合的交叉注意力模块;(2)针对大规模数据集优化的预处理流程;(3)鲁棒的数据增强技术。在大规模灾害数据集上的实验表明,整体分类准确率达94.90%。模型能有效区分不同损毁等级,对数据缺失具有较强鲁棒性。该系统显著提升评估速度与准确性,有助于应急人员优先制定干预措施。本工作通过融合多时相影像与深度学习,推动自动化灾损检测发展,提供可扩展的实时响应方案。
原文摘要 · Abstract (English)
Timely and accurate disaster damage assessment is crucial for effective emergency response, resource allocation, and recovery. Traditional methods, which often rely on manual inspections or sparse data, are typically slow and error-prone. This paper introduces a novel framework leveraging remote sensing imagery and deep learning to automate building damage classification. Using pre- and post-disaster satellite imagery, our model categorizes buildings into four damage levels: no damage, minor damage, major damage, and destroyed. The core innovation is a multi-modal attention mechanism that fuses bi-temporal features to explicitly detect and assess structural changes. We employ a lightweight ConvNeXT-Tiny backbone to ensure efficient processing without compromising performance. Key contributions include: (1) a cross-attention module for multi-modal data fusion, (2) an optimized preprocessing pipeline for large-scale datasets, and (3) robust data augmentation techniques. Experiments on a large-scale disaster dataset demonstrate an overall classification accuracy of 94.90%. The model effectively discriminates between damage categories and remains resilient to incomplete data. This system significantly improves assessment speed and accuracy, aiding emergency responders in prioritizing interventions. This work advances automated disaster damage detection by integrating multi-temporal imagery with deep learning, offering a scalable solution for real-time response.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。