用标注者身份信息训练大模型,提升仇恨言论检测的公平性。
Algorithmic Fairness in NLP: Persona-Infused LLMs for Human-Centric Hate Speech Detection
- 通过浅层提示和RAG构建深度身份画像,注入标注者背景。
- 同群体标注者身份使模型对特定群体更敏感,降低误判率。
- 为构建更公平的仇恨言论检测系统提供心理与技术结合的新思路。
本文研究将标注者身份信息融入大语言模型(Persona-LLMs)对仇恨言论检测敏感性的影响,尤其关注标注者与目标群体身份相同或不同所引发的偏见。我们采用谷歌Gemini和OpenAI GPT-4.1-mini模型,结合浅层身份提示与基于检索增强生成(RAG)的深度上下文身份构建方法,引入更丰富的身份特征。分析了同群组(in-group)与异群组(out-group)标注者身份对模型检测性能及公平性的影响。研究融合心理学中群体身份理论与先进自然语言处理技术,表明在大模型中融入社会人口属性可缓解自动化仇恨言论检测中的偏见。结果揭示了基于身份的方法在减少偏见方面的潜力与局限,为开发更公正的检测系统提供了关键洞见。
原文摘要 · Abstract (English)
In this paper, we investigate how personalising Large Language Models (Persona-LLMs) with annotator personas affects their sensitivity to hate speech, particularly regarding biases linked to shared or differing identities between annotators and targets. To this end, we employ Google's Gemini and OpenAI's GPT-4.1-mini models and two persona-prompting methods: shallow persona prompting and a deeply contextualised persona development based on Retrieval-Augmented Generation (RAG) to incorporate richer persona profiles. We analyse the impact of using in-group and out-group annotator personas on the models' detection performance and fairness across diverse social groups. This work bridges psychological insights on group identity with advanced NLP techniques, demonstrating that incorporating socio-demographic attributes into LLMs can address bias in automated hate speech detection. Our results highlight both the potential and limitations of persona-based approaches in reducing bias, offering valuable insights for developing more equitable hate speech detection systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。