分析标注者与目标的背景如何影响仇恨言论标注中的偏见。
Human and LLM Biases in Hate Speech Annotations: A Socio-Demographic Analysis of Annotators and Targets
- 对比标注者与目标的社会属性,揭示偏见来源。
- 发现人类标注者存在显著偏见,且强度与普遍性各异。
- 对比大模型标注偏见,发现其与人类不同,适合优化AI系统设计。
在线平台的兴起加剧了仇恨言论的传播,亟需可扩展且高效的检测手段。然而,仇恨言论检测系统的准确性高度依赖人工标注数据,而这类数据天然易受偏见影响。尽管已有研究关注此问题,但标注者特征与被标注目标特征之间的相互作用仍不清楚。本文利用包含丰富社会人口学信息的大型数据集,深入分析标注者与目标属性如何共同影响标注偏见。研究发现广泛存在的偏见,并定量描述其强度与分布特征,揭示显著差异。此外,我们还对比了基于角色设定的大语言模型(persona-based LLMs)的偏见表现,结果表明:虽然大模型也表现出偏见,但其模式与人类标注者显著不同。本研究为仇恨言论标注中的偏见提供了更细致的理解,并为人工智能驱动的检测系统设计带来新洞见。
原文摘要 · Abstract (English)
The rise of online platforms exacerbated the spread of hate speech, demanding scalable and effective detection. However, the accuracy of hate speech detection systems heavily relies on human-labeled data, which is inherently susceptible to biases. While previous work has examined the issue, the interplay between the characteristics of the annotator and those of the target of the hate are still unexplored. We fill this gap by leveraging an extensive dataset with rich socio-demographic information of both annotators and targets, uncovering how human biases manifest in relation to the target's attributes. Our analysis surfaces the presence of widespread biases, which we quantitatively describe and characterize based on their intensity and prevalence, revealing marked differences. Furthermore, we compare human biases with those exhibited by persona-based LLMs. Our findings indicate that while persona-based LLMs do exhibit biases, these differ significantly from those of human annotators. Overall, our work offers new and nuanced results on human biases in hate speech annotations, as well as fresh insights into the design of AI-driven hate speech detection systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。