arXiv:2605.01168cs.CL2026-05中稿 · the 5th Workshop o…被引 2

研究如何预测人工标注中对不当语言的分歧程度。

Quantifying and Predicting Disagreement in Graded Human Ratings

论文配图:Quantifying and Predicting Disagreement in Graded Human Ratings
图 1 · 摘自论文原文
  • 用文本特征预测标注者分歧水平,提出对立指数衡量观点冲突。
  • 模型预测分歧方差与实际结果有中度正相关,效果稳定。
  • 观点对立强的样本更难预测,现有模型常低估其难度。

人类标注者在标注任务中常存在分歧,且不同内容引发的分歧程度不同。本文研究了对不当语言(如辱骂性语言、仇恨言论、有毒语言)进行分级标注时的标注变异模式,探讨能否通过文本特征预测分歧程度。我们提出对立指数(Opposition Index),量化特定内容上标注者之间的观点对立程度,并研究高对立样本的可预测性。结果表明,预测的分歧方差与实际观测值之间存在中度正相关。两种方法表现相当:直接预测方差值,或从预测的标注分布估算方差。在对立视角预测方面,高对立指数的样本更难预测,且模型普遍低估其复杂性。

原文摘要 · Abstract (English)

It is increasingly recognized that human annotators do not always agree, and such disagreement is inherent in many annotation tasks. However, not all instances in a given task elicit the same degree of opinion divergence. In this paper, we investigate annotation variation patterns in graded human ratings for inappropriate languages, including offensive language, hate speech, and toxic language perception. We examine whether the degree of annotation disagreement can be predicted from textual features. We further propose the Opposition Index, a metric that quantifies perspective opposition among annotators on a given item, and investigate the predictability of instances with potentially opposing human opinions. Our results show a moderate positive correlation between estimated and observed annotation variance. We find that two approaches achieve comparable performance in variance prediction: directly predicting the variance value and estimating it from predicted annotation distributions. Our results on opposition perspective prediction show that items with high opposition index values are more difficult to predict and are often underestimated by models.

标注分歧文本评估语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。