用虚拟身份评估大模型对仇恨言论的跨群体理解能力
From Self to Other: Evaluating Demographic Perspective-Taking in LLM Hate Speech Annotation
- 让大模型扮演不同身份,模拟各群体对仇恨言论的判断差异
- 只有特定配置的Llama 3.1在跨群体一致性上表现最佳
- 适合用于需要反映真实社会分歧的自动内容审核场景
仇恨言论判定具有高度主观性:不同人群对同一内容的感受截然不同。收集多群体标注数据成本高且难以扩展。角色化大语言模型(通过提示设定特定身份)被提出用于规模化模拟多元视角。我们评估了人类社会判断的三个维度:(i) 不同群体角色是否以类人方式产生分歧(跨群体分歧);(ii) 模型是否对自身所属群体的攻击更敏感(本群敏感性);(iii) 是否能准确预测其他群体的反应(共情预测)。结果表明,无一模型能稳定覆盖全部三维度,性能高度依赖模型本身,仅靠简单身份提示无法可靠生成分歧模式。但使用Llama 3.1进行共情提示时,在多数人口维度上达到最高跨群体一致性和最接近人类分歧模式的整体逼近,表明该配置或可作为更贴近人类判断的自动化标注方案。
原文摘要 · Abstract (English)
Hate speech detection is inherently subjective: people from different demographic groups perceive the same content very differently. Collecting enough annotations from multiple demographic groups is costly and difficult to scale. Persona-conditioned Large Language Models (models prompted to adopt a specific demographic identity) have been proposed as a way to simulate diverse perspectives at scale. But do they actually reflect how different groups disagree? We evaluate three aspects of human social judgement: (i) whether personas from different groups disagree in human-like ways (inter-group disagreement), (ii) whether they become more sensitive when content targets their own identity (in-group sensitivity), and (iii) whether they can accurately predict how another group would react (vicarious prediction). Our results show that no model consistently captures all three dimensions, and performance is highly model-dependent and does not emerge reliably from minimal identity prompts alone. However, vicarious prompting with Llama 3.1 yields the highest cross-group agreement in most demographic axes and provides the closest overall approximation to human disagreement patterns, indicating that this configuration may provide a more reliable setting for automatic annotation aligned with human judgements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。