对比无人机与卫星影像的灾后建筑损毁标注,发现超四成不一致,影响AI评估可靠性。
Now you see it, Now you don't: Damage Label Agreement in Drone & Satellite Post-Disaster Imagery
- 统一标注标准和建筑位置,对比三场飓风中1.5万栋建筑的损毁标签。
- 卫星标注比无人机低20.43%以上,且两类数据分布差异极显著(p<5e-175)。
- 提醒用AI做灾损评估时需警惕数据源偏差,避免误导救援决策。
本文对飓风伊恩、迈克尔与哈维期间,15,814栋建筑的卫星与无人机航拍影像中的损毁标注进行审计,发现29.02%的标注存在分歧,且两类数据来源的损毁分布显著不同。当前尚无研究系统比较无人机与卫星影像在建筑损毁评估中的标注一致性。此前相关工作受限于标注体系不一、建筑位置错位及数据量不足。本研究通过统一标注标准与建筑位置,覆盖三场飓风,样本量为最相关先前研究的19.05倍。分析显示,卫星标注至少低估损毁20.43%(p<1.2×10^-117),且卫星与无人机标注分布差异极显著(p<5.1×10^-175)。这表明基于任一数据源训练的计算机视觉与机器学习模型无法准确反映真实场景,可能造成误判,带来伦理风险与社会危害。为此,论文提出四项建议,以提升实际部署中CV/ML灾损评估系统的可靠性与透明度。
原文摘要 · Abstract (English)
This paper audits damage labels derived from coincident satellite and drone aerial imagery for 15,814 buildings across Hurricanes Ian, Michael, and Harvey, finding 29.02% label disagreement and significantly different distributions between the two sources, which presents risks and potential harms during the deployment of machine learning damage assessment systems. Currently, there is no known study of label agreement between drone and satellite imagery for building damage assessment. The only prior work that could be used to infer if such imagery-derived labels agree is limited by differing damage label schemas, misaligned building locations, and low data quantities. This work overcomes these limitations by comparing damage labels using the same damage label schemas and building locations from three hurricanes, with the 15,814 buildings representing 19.05 times more buildings considered than the most relevant prior work. The analysis finds satellite-derived labels significantly under-report damage by at least 20.43% compared to drone-derived labels (p<1.2x10^-117), and satellite- and drone-derived labels represent significantly different distributions (p<5.1x10^-175). This indicates that computer vision and machine learning (CV/ML) models trained on at least one of these distributions will misrepresent actual conditions, as the differing satellite and drone-derived distributions cannot simultaneously represent the distribution of actual conditions in a scene. This potential misrepresentation poses ethical risks and potential societal harm if not managed. To reduce the risk of future societal harms, this paper offers four recommendations to improve reliability and transparency to decisio-makers when deploying CV/ML damage assessment systems in practice
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。