提出新方法精准定位标注者群体间观点分歧根源,提升仇恨言论检测可靠性。
Are we chasing ghosts? Quantifying unattributable polarization, and attributing the rest to annotator groups
- 设计可检验显著性的极化归因度量,避免虚假信号干扰。
- 实测每条评论仅需20名标注者即可稳定估计极化程度。
- 发现性别与种族是极化现象的关键解释因素,群体差异随距离增大而增强。
标准一致性度量常无法捕捉少数群体与多数群体之间的系统性意见差异,影响仇恨言论与毒性检测等任务。极化被提出作为更稳健的区分方式,但现有方法缺乏对极化来源的可操作归因工具。我们评估现有方法,识别出两个现实场景中的主要局限:(1) 无法归因于任何已知或潜在群体的‘固有’极化存在;(2) 不同方向极化在聚合标注中相互抵消。为此,我们提出一种新度量,能统计检验标注者群体极化归因的显著性,同时规避上述问题,并开源了Python实现库。实验表明,每条评论仅需20名标注者即可获得可靠估计。在四个主观NLP数据集上应用该方法,发现性别与种族始终能解释极化模式,且群体间差异随分组距离增加而增强。
原文摘要 · Abstract (English)
Standard agreement metrics often fail to capture systematic differences in opinion between minority and majority-group annotators, jeopardizing tasks such as hate speech and toxicity detection. Polarization has recently been proposed as a more robust way of distinguishing minor disagreements from systematic differences in opinion, but existing approaches do not provide practical tools for attributing it to specific annotator groups. We evaluate current methods and identify two major limitations in realistic settings: (1) the presence of ``inherent'' polarization that cannot be attributed to any known or latent groups, and (2) opposing polarization effects canceling each other out in aggregated annotations. To address these issues, we introduce a new metric that measures and tests the statistical significance of polarization attribution for annotator groups while avoiding these limitations, as well as an open-source Python library implementation, finding that no more than 20 annotators are needed per comment for reliable estimation. We apply our method to four subjective NLP datasets and find that gender and race consistently explain polarization patterns, while differences between annotator groups become stronger as the groups are further apart.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。