多数投票掩盖争议边界分歧,导致模型误判仇恨言论临界点。
Majority Vote Silences Minority Values: Annotator Disagreement at the Hate/Offensive Boundary in HateXplain

- 用多数投票合并标注分歧,却在仇恨与不当言论交界处集中出现42.6%分歧。
- 模型在标注不一致的文本上准确率从80%降至58%,且软标签模型更差。
- 当前评估无法发现模型对边界案例的错误高自信,需从标注设计源头解决。
仇恨言论标注流程常将标注者分歧统一为多数投票标签后用于训练。我们发现这种聚合并非中立:HateXplain数据集中42.6%的标注分歧集中在仇恨/不当言论的边界,表明标注者对仇恨起始阈值存在差异(卡方检验 = 135.199,df = 2,p < 0.0001)。无论是硬标签BERT模型(模型A)还是软标签模型(模型B),在标注一致文本上的准确率约为80%,而在分歧文本上均下降至约58%,差异显著(p < 0.0001)。每标注者多头模型(模型C)进一步拉大差距至28个百分点,且将不当言论分歧的准确率压低至0.245。关键的是,模型A在边界错误案例上表现出更高置信度(0.710 vs. 0.495,p < 0.0001),导致标准评估指标无法识别其失效。三种逐步复杂的下游干预均未能恢复边界准确率。我们认为问题本质是结构性的:多数投票将有争议的判断呈现为真实标签,模型继承了这种虚假确定性。解决之道必须在标注设计阶段进行上游干预。
原文摘要 · Abstract (English)
Hate speech annotation pipelines routinely collapse annotator disagreement into majority vote labels before training. We show that this aggregation is not neutral: 42.6% of all annotator disagreement in HateXplain concentrates specifically at the hate/offensive boundary, a pattern consistent with annotators applying different thresholds for where hate begins (chi-squared = 135.199, df = 2, p < 0.0001). Both a hard-label BERT model (Model A) and a soft-label model (Model B) drop 22 percentage points in accuracy from agreed posts (~80%) to disagreement posts (~58%), confirmed at p < 0.0001. A per-annotator multi-head model (Model C) widens this gap further to 28 points while collapsing offensive disagreement accuracy to 0.245. Critically, Model A expresses significantly higher confidence on boundary case errors than Model C (0.710 vs. 0.495, p < 0.0001), meaning standard evaluation metrics will not detect the failure. Three downstream interventions of increasing sophistication all fail to recover boundary accuracy. We argue the problem is structural. Majority vote presents a contested judgment as ground truth, and models inherit that false certainty. The intervention must be upstream in annotation design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。