挑战数据标注中'真实标签'的神话,揭示人类分歧是重要信号而非噪音。
The Consensus Trap: Dissecting Subjectivity and the "Ground Truth" Illusion in Data Annotation
- 分析346篇论文,指出标注系统因位置不透明和模型中介导致偏见固化。
- 发现地理霸权使西方标准被当作普适基准,标注员为生存被迫迎合请求者。
- 主张将标注目标从寻找唯一正确答案转向记录多元人类经验。
在机器学习中,'真实标签'指用于训练和评估模型的假设正确标签。然而,这一基础范式基于一种实证谬误,将人类分歧视为技术噪声而非关键的社会技术信号。本系统文献综述分析了2020至2025年间在ACL、AIES、CHI、CSCW、EAAMO、FAccT和NeurIPS七个顶级会议发表的研究,探讨数据标注实践中促成'共识陷阱'的机制。对346篇论文的反思性主题分析表明,位置可见性缺失与近期向'人作为验证者'模型的架构转变——特别是对模型中介标注的依赖——引入了深层锚定偏差,并实质上将人类声音排除在流程之外。我们进一步揭示地理霸权如何将西方规范强加为普遍标准,常由处境脆弱的标注员通过讨好请求者以避免经济惩罚来执行。批判'噪声传感器'谬误(即统计模型误将多元性视为错误),主张重拾分歧作为构建文化适应性模型所必需的高保真信号。为此,我们提出一个多元化标注基础设施的路线图,将目标从发现单一'正确'答案转向描绘人类经验的多样性。
原文摘要 · Abstract (English)
In machine learning, "ground truth" refers to the assumed correct labels used to train and evaluate models. However, the foundational "ground truth" paradigm rests on a positivistic fallacy that treats human disagreement as technical noise rather than a vital sociotechnical signal. This systematic literature review analyzes research published between 2020 and 2025 across seven premier venues: ACL, AIES, CHI, CSCW, EAAMO, FAccT, and NeurIPS, investigating the mechanisms in data annotation practices that facilitate this "consensus trap". Our reflexive thematic analysis of 346 papers reveals that systemic failures in positional legibility, combined with the recent architectural shift toward human-as-verifier models, specifically the reliance on model-mediated annotations, introduce deep-seated anchoring bias and effectively remove human voices from the loop. We further demonstrate how geographic hegemony imposes Western norms as universal benchmarks, often enforced by the performative alignment of precarious data workers who prioritize requester compliance over honest subjectivity to avoid economic penalties. Critiquing the "noisy sensor" fallacy, where statistical models misdiagnose pluralism as error, we argue for reclaiming disagreement as a high-fidelity signal essential for building culturally competent models. To address these systemic tensions, we propose a roadmap for pluralistic annotation infrastructures that shift the objective from discovering a singular "right" answer to mapping the diversity of human experience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。