用多模态分歧仲裁提升灾后街景图像损毁评估的准确率与可信度。
DamageArbiter: A Multimodal Arbitration Framework for Disaster Damage Assessment from Street-View Imagery

- 通过图像与文本模型分歧驱动,用轻量逻辑回归仲裁决策。
- 准确率达75.85%,过自信错误从70.58%降至16.45%。
- 适合需要高可靠性评估的应急响应与灾后决策场景。
利用计算机视觉分析街景图像可实现快速、超本地化的灾后损毁评估,但现有方法多依赖黑箱预训练视觉模型,缺乏可解释性与可靠性。本文提出DamageArbiter,一种基于多模态分歧驱动的仲裁框架,融合单模态与多模态模型优势,采用轻量逻辑回归元分类器处理预测分歧。基于2,556张灾后街景图像(配有人工或大语言模型生成的文本描述),系统对比了DamageArbiter与微调的单模态(仅图像/仅文本)模型及CLIP类多模态模型在分类性能与过自信误差上的表现。结果表明,DamageArbiter将准确率提升至75.85%,马修斯相关系数(MCC)达0.6188,显著优于最佳基线(文本模型:63.07%准确率,0.4126 MCC;图像模型:74.33%准确率,0.5947 MCC;CLIP:74.22%准确率,0.5915 MCC)。过自信分析显示,其将过自信误差从最优基线(图像仅模型)的70.58%大幅降至16.45%。研究强调,仅看准确率不足以评估灾损分类模型,必须结合过自信误差衡量模型可靠性。DamageArbiter为街景图像驱动的快速灾损评估提供了更可靠的框架。
原文摘要 · Abstract (English)
Analyzing street-view imagery with computer vision models offers a promising approach for rapid, hyperlocal disaster damage assessment, but existing approaches typically rely on black-box pre-trained vision models, which lack interpretability and reliability. This study proposes DamageArbiter, a multimodal disagreement-driven arbitration framework designed to improve the accuracy and reliability of street-view-based damage assessment. DamageArbiter leverages the complementary strengths of unimodal and multimodal models and employs a lightweight logistic regression meta-classifier to arbitrate cases in which model predictions disagree. Using 2,556 post-disaster street-view images, paired with manually generated or large language model (LLM)-generated text descriptions, we systematically compared DamageArbiter with fine-tuned unimodal (image-only and text-only) models and CLIP-based multimodal models in terms of classification performance and overconfidence errors. Results show that DamageArbiter improved accuracy to 75.85% and the Matthews correlation coefficient (MCC) to 0.6188, compared with the best-performing text-only baseline (63.07% accuracy, 0.4126 MCC), image-only baseline (74.33% accuracy, 0.5947 MCC), and CLIP baseline (74.22% accuracy, 0.5915 MCC). The overconfidence analysis further reveals that DamageArbiter substantially reduced the overconfidence error from 70.58% for the best-performing baseline, the image-only ViT model, to 16.45%. Overall, this study demonstrates that accuracy alone is insufficient for evaluating disaster damage classification models and highlights the importance of measuring overconfidence errors as part of model reliability assessment. DamageArbiter thus offers a more reliable framework for rapid, hyperlocal disaster damage assessment from street-view imagery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。