arXiv:2512.20577cs.LG2025-12被引 1

用统计方法评估标注数据质量,提升机器学习训练效果。

Improving ML Training Data with Gold-Standard Quality Metrics

  • 通过多次标注迭代测量一致性,识别高质量数据。
  • 减少重复标注即可获得高质数据,无需每项都多人标注。
  • 标注员适应期不足以降低错误,需持续监控标注表现。

人工标注数据对众多机器学习任务至关重要,但数据质量控制在文献中长期被忽视,因标注质量随标注过程变化显著。本文提出基于统计方法的评估与提升手标签训练数据质量的策略,通过测量标注一致性和共识度来实现。研究发现,若在多次标注迭代中记录共识度指标,其方差下降可作为数据质量提升的可靠信号。此外,本文展示了一种无需为每个样本进行多轮标注即可获取高质量数据的方法,并指出仅依靠标注员适应期不足以有效降低标注错误。

原文摘要 · Abstract (English)

Hand-tagged training data is essential to many machine learning tasks. However, training data quality control has received little attention in the literature, despite data quality varying considerably with the tagging exercise. We propose methods to evaluate and enhance the quality of hand-tagged training data using statistical approaches to measure tagging consistency and agreement. We show that agreement metrics give more reliable results if recorded over multiple iterations of tagging, where declining variance in such recordings is an indicator of increasing data quality. We also show one way a tagging project can collect high-quality training data without requiring multiple tags for every work item, and that a tagger burn-in period may not be sufficient for minimizing tagger errors.

数据质量标注评估机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。