arXiv:2607.15455stat.MLcs.AI2026-07

用部分专家审核的噪声标签,纠正自动化标签偏差。

Design-Based Supervised Learning with Noisy Human Labels

论文配图:Design-Based Supervised Learning with Noisy Human Labels
图 1 · 摘自论文原文
  • 基于抽样审计设计,利用专家审核数据修正噪声人类标签
  • 在合成与维基百科实验中,误差降低10%-17%,覆盖率达标
  • 适合处理带噪声标签的自动化数据标注场景

研究人员越来越多地使用自动化分类器对非结构化数据进行标注以支持统计分析。现有校正方法通常依赖概率抽样的审计集来纠正自动化标签错误,但假设审计标签为正确。实际上,人工审计标签常含噪声,且仅部分样本经专家审定。本文提出部分审定的基于设计的监督学习(PA-DSL),利用专家审定案例修正噪声人类标签,并用修正后的审计信息对全量自动化标签的分析结果去偏。当审计与审定概率已知时,该估计器适用于广泛下游分析。在合成数据和Wikipedia Detox半合成实验中,当噪声人类标签包含可恢复信号时,PA-DSL保持名义覆盖,相对仅使用审定标签,均方根误差降低10%-17%。

原文摘要 · Abstract (English)

Researchers increasingly use automated classifiers to label unstructured data for statistical analysis. Existing rectification methods can correct errors in these automated labels using a probability-sampled audit set, but they usually treat the audit labels as correct. In practice, human audit labels are often noisy, and only some audited items are reviewed by an expert or adjudicator. We propose Partially Adjudicated Design-Based Supervised Learning (PA-DSL), a method for this setting. It uses adjudicated cases to correct noisy human labels and then uses the corrected audit information to debias analyses based on the full set of automated labels. The estimator is valid for a broad class of downstream analyses when the audit and adjudication probabilities are known. In synthetic and Wikipedia Detox semi-synthetic experiments, PA-DSL maintains nominal coverage and reduces RMSE by 10-17% relative to using only adjudicated labels when noisy human labels contain recoverable signal.

标签纠错噪声数据审计设计统计推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。