多源航拍影像标注质量差异大,低分辨率影像需更多人工审核。
Looks Can be Deceiving: Annotator and Reviewer Performance Across Imagery Sources in Crowd-Sourced Aerial Damage Assessment

- 对比无人机、有人机、卫星三类影像的标注一致性
- 卫星影像标注修改率高达36.95%,是最高者
- 建议按分辨率动态分配审核资源,提升数据质量
本文首次实证研究了多源遥感影像中标注者与审核者的表现,评估了在无人机、有人驾驶航空和卫星图像中的灾后建筑损毁标注。现有数据集多依赖单一来源,缺乏高效的人工标注资源配置标准。本研究基于9次灾害的损毁评估数据集,包含20041栋无人机影像建筑、20695栋有人机影像建筑和33392栋卫星影像建筑,由187名标注者完成标注,并经过两轮质检:单人审核和共识委员会复审。分析发现,初始标注在不同分辨率影像中被修订比例呈递增趋势(有人机25.27%,卫星36.95%),且每阶段均如此。即使经单人审核,仍存在显著分歧:最终委员会仍修正了6.85%无人机、14.05%有人机、20.86%卫星标注。表明均匀分配审核任务会留下大量低分辨率影像的误差。结合自适应任务分配与预算感知质量控制,本文提出三项改进建议。
原文摘要 · Abstract (English)
This paper presents the first known empirical investigation of annotator and reviewer performance across multi-source remotely sensed imagery, evaluating human labeling across drone, crewed aviation, and satellite views. Because existing aerial imagery datasets rely predominantly on single-source imagery, there is no currently established state of practice for efficiently allocating human labor to curate large-scale, multi-source aerial datasets. This work addresses this limitation by analyzing annotator and reviewer performance within a post-disaster building damage assessment dataset of 9 disasters, where 20041 buildings in drone, 20695 buildings in crewed aviation, and 33392 buildings in satellite imagery were labeled. These labels, provided by 187 annotators, were then refined through two successive quality-control stages: a single-reviewer pass followed by a consensus-committee review. Our analysis reveals two findings that raise questions for standard crowd-sourcing practices. First, initial annotations were revised by the final committee at rates that rise steeply from higher- to lower-resolution sources (25.27% for crewed aviation and 36.95% for satellite), with the same ordering at every observed workflow stage. Second, a single individual review reduced but did not resolve this disagreement: after review, the committee still revised 6.85% of drone, 14.05% of crewed, and 20.86% of satellite labels. These observations suggest that, in workflows like this one, uniform review allocation leaves the most residual disagreement in lower-resolution imagery. Based on this evidence, and consistent with prior work on adaptive task assignment and budget-aware quality control, this paper offers three recommendations for multi-source dataset curation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。