用机器预测人工标注,既省钱又保质,还能自动识别难样本。
No Need to Sacrifice Data Quality for Quantity: Crowd-Informed Machine Annotation for Cost-Effective Understanding of Visual Data
- 用卷积神经网络预测人群回答,实现高效自动化标注
- 在汽车数据集上节省超20%成本,准确识别人类不确定样本
- 预测的软标签后验分布可作后续推理先验,减少人工标注需求
视觉数据标注成本高昂且耗时。众包系统虽能并行化标注,但仍有限制。本文提出一种框架,可在不牺牲结果可靠性的情况下大规模进行视觉数据质量检查。通过向标注者提出带离散答案的简单问题,利用训练好的卷积神经网络自动预测人群响应。不同于以往直接预测软标签的方法,本方法以每任务的软标签后验分布为训练目标,引入狄利克雷先验以保证解析可得性。我们在两个真实的汽车自动驾驶数据集上验证该方法,结果显示模型可完全自动化大量任务,使成本降低达两位数百分比以上。模型可靠预测人类不确定性,有助于精准筛选难点样本。此外,模型输出的软标签后验分布可作为后续推断的先验,显著减少对人工标注者的依赖,进一步降低成本并提升人力使用效率。
原文摘要 · Abstract (English)
Labeling visual data is expensive and time-consuming. Crowdsourcing systems promise to enable highly parallelizable annotations through the participation of monetarily or otherwise motivated workers, but even this approach has its limits. The solution: replace manual work with machine work. But how reliable are machine annotators? Sacrificing data quality for high throughput cannot be acceptable, especially in safety-critical applications such as autonomous driving. In this paper, we present a framework that enables quality checking of visual data at large scales without sacrificing the reliability of the results. We ask annotators simple questions with discrete answers, which can be highly automated using a convolutional neural network trained to predict crowd responses. Unlike the methods of previous work, which aim to directly predict soft labels to address human uncertainty, we use per-task posterior distributions over soft labels as our training objective, leveraging a Dirichlet prior for analytical accessibility. We demonstrate our approach on two challenging real-world automotive datasets, showing that our model can fully automate a significant portion of tasks, saving costs in the high double-digit percentage range. Our model reliably predicts human uncertainty, allowing for more accurate inspection and filtering of difficult examples. Additionally, we show that the posterior distributions over soft labels predicted by our model can be used as priors in further inference processes, reducing the need for numerous human labelers to approximate true soft labels accurately. This results in further cost reductions and more efficient use of human resources in the annotation process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。