研究大模型在仇恨内容识别中的偏见,发现其对不同人群和地域敏感。
Hate Personified: Investigating the role of LLMs in content moderation
- 通过人物角色和地理线索增强提示,测试模型对多元群体的响应差异。
- 模型对数字锚点敏感,能反映社区举报与对立观点的影响。
- 结果表明应谨慎使用大模型做跨文化内容审核,需考虑群体差异。
对于仇恨检测这类主观任务,不同人群对仇恨的理解存在差异,而大语言模型(LLM)能否体现多元群体需求尚不明确。本文通过在提示中引入额外上下文,系统分析了LLM对地理先验、人物属性及数值信息的敏感性,以评估其对不同群体需求的反映程度。实验覆盖两个LLM、五种语言和六个数据集,结果表明:模拟人物属性会引发标注变异;引入地理信号可提升区域一致性;模型对数值锚点敏感,说明其能吸收社区举报行为与对立观点的影响。本研究提供初步应用指南,揭示了在文化敏感场景下使用LLM进行内容审核的复杂性。
原文摘要 · Abstract (English)
For subjective tasks such as hate detection, where people perceive hate differently, the Large Language Model's (LLM) ability to represent diverse groups is unclear. By including additional context in prompts, we comprehensively analyze LLM's sensitivity to geographical priming, persona attributes, and numerical information to assess how well the needs of various groups are reflected. Our findings on two LLMs, five languages, and six datasets reveal that mimicking persona-based attributes leads to annotation variability. Meanwhile, incorporating geographical signals leads to better regional alignment. We also find that the LLMs are sensitive to numerical anchors, indicating the ability to leverage community-based flagging efforts and exposure to adversaries. Our work provides preliminary guidelines and highlights the nuances of applying LLMs in culturally sensitive cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。