用多尺度注意力提升无人机灾后损毁识别精度
MCANet: A Multi-Scale Class-Specific Attention Network for Multi-Label Post-Hurricane Damage Assessment using UAV Imagery
- 设计多头类别特定注意力模块,捕捉不同尺度空间特征
- 在4494张飓风后航拍图上达到91.75% mAP,部分难类提升超6%
- 适合灾害应急、智能救援与数字孪生系统应用
快速准确的灾后损毁评估对救灾恢复至关重要。现有基于CNN的方法难以捕捉多尺度空间特征,且难以区分视觉相似或共现的损毁类型。为此,我们提出MCANet,一种多标签分类框架,能学习多尺度表征并为每类损毁自适应关注相关区域。MCANet采用基于Res2Net的分层主干网络以丰富跨尺度空间上下文,并引入多头类别特定残差注意力模块增强判别能力。每个注意力分支聚焦不同空间粒度,平衡局部细节与全局上下文。我们在飓风迈克尔后采集的4,494张无人机图像组成的RescueNet数据集上评估,MCANet实现91.75%的平均精度(mAP),优于ResNet、Res2Net、VGG、MobileNet、EfficientNet和ViT。使用八个注意力头时性能进一步提升至92.35%,对“道路阻塞”等困难类别的平均精度提升超过6%。类激活图验证了MCANet定位损毁相关区域的能力,支持可解释性。其输出可用于灾后风险制图、应急路径规划及数字孪生灾备系统。未来工作可结合灾害知识图谱与多模态大语言模型,提升对未见灾害的适应性与语义理解能力。
原文摘要 · Abstract (English)
Rapid and accurate post-hurricane damage assessment is vital for disaster response and recovery. Yet existing CNN-based methods struggle to capture multi-scale spatial features and to distinguish visually similar or co-occurring damage types. To address these issues, we propose MCANet, a multi-label classification framework that learns multi-scale representations and adaptively attends to spatially relevant regions for each damage category. MCANet employs a Res2Net-based hierarchical backbone to enrich spatial context across scales and a multi-head class-specific residual attention module to enhance discrimination. Each attention branch focuses on different spatial granularities, balancing local detail with global context. We evaluate MCANet on the RescueNet dataset of 4,494 UAV images collected after Hurricane Michael. MCANet achieves a mean average precision (mAP) of 91.75%, outperforming ResNet, Res2Net, VGG, MobileNet, EfficientNet, and ViT. With eight attention heads, performance further improves to 92.35%, boosting average precision for challenging classes such as Road Blocked by over 6%. Class activation mapping confirms MCANet's ability to localize damage-relevant regions, supporting interpretability. Outputs from MCANet can inform post-disaster risk mapping, emergency routing, and digital twin-based disaster response. Future work could integrate disaster-specific knowledge graphs and multimodal large language models to improve adaptability to unseen disasters and enrich semantic understanding for real-world decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。