arXiv:2507.07216cs.LGcs.AI2025-07

针对标签偏见导致的错误标注,提出可解耦的可信学习框架。

Bias-Aware Mislabeling Detection via Decoupled Confident Learning

  • 分离标签置信度与偏见风险,识别受社会群体影响的标注错误。
  • 在仇恨言论检测中表现优于现有方法,提升数据可靠性。
  • 适合关注数据质量与公平性的组织使用。

可靠数据是现代组织系统的基础。标签偏见——即标签中系统性错误,其质量在不同社会群体间存在差异——是量化分析中的核心挑战。此类偏见在多个关键领域被广泛证实且亟需解决,但有效应对方法仍匮乏。本文提出解耦可信学习(DeCoLe),一种基于机器学习的框架,专门用于检测受标签偏见影响的数据集中错误标注实例,实现偏见感知的误标检测,助力数据质量提升。我们从理论上证明了DeCoLe的有效性,并在仇恨言论检测这一标签偏见问题突出的领域进行评估。实证结果表明,DeCoLe在偏见感知的误标检测上持续优于其他方法。本工作解决了偏见感知误标检测的挑战,为将DeCoLe融入组织数据管理实践提供了指导,以增强数据可靠性。

原文摘要 · Abstract (English)

Reliable data is a cornerstone of modern organizational systems. A notable data integrity challenge stems from label bias, which refers to systematic errors in a label, a covariate that is central to a quantitative analysis, such that its quality differs across social groups. This type of bias has been conceptually and empirically explored and is widely recognized as a pressing issue across critical domains. However, effective methodologies for addressing it remain scarce. In this work, we propose Decoupled Confident Learning (DeCoLe), a principled machine learning based framework specifically designed to detect mislabeled instances in datasets affected by label bias, enabling bias aware mislabelling detection and facilitating data quality improvement. We theoretically justify the effectiveness of DeCoLe and evaluate its performance in the impactful context of hate speech detection, a domain where label bias is a well documented challenge. Empirical results demonstrate that DeCoLe excels at bias aware mislabeling detection, consistently outperforming alternative approaches for label error detection. Our work identifies and addresses the challenge of bias aware mislabeling detection and offers guidance on how DeCoLe can be integrated into organizational data management practices as a powerful tool to enhance data reliability.

数据质量偏见检测可信学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。