arXiv:2506.08593cs.CL2025-06被引 1

研究人格特质如何影响大模型识别仇恨言论,发现不同人格提示导致判断差异。

Hateful Person or Hateful Model? Investigating the Role of Personas in Hate Speech Detection by Large Language Models

  • 用MBTI人格提示引导大模型进行仇恨言论分类
  • 模型判断结果随人格提示显著变化,与真实标签不一致
  • 提醒在人机协作标注中需谨慎设计人格提示

仇恨言论检测是一项社会敏感且主观性较强的任务,判断常受个人特质影响。尽管已有研究探讨社会人口学因素对标注的影响,但人格特质对大语言模型(LLMs)的影响仍基本未被探索。本文首次系统研究了人格提示在仇恨言论分类中的作用,聚焦于MBTI人格类型。人类标注调查证实MBTI维度显著影响标注行为。将此扩展至大模型,我们使用四种开源模型,通过MBTI人格提示在三个仇恨言论数据集上进行评估。分析揭示了显著的人格驱动差异,包括与真实标签的不一致、不同人格间的分歧以及逻辑输出层面的偏差。这些发现强调,在基于大模型的标注流程中,需审慎定义人格提示,以确保公平性并契合人类价值观。

原文摘要 · Abstract (English)

Hate speech detection is a socially sensitive and inherently subjective task, with judgments often varying based on personal traits. While prior work has examined how socio-demographic factors influence annotation, the impact of personality traits on Large Language Models (LLMs) remains largely unexplored. In this paper, we present the first comprehensive study on the role of persona prompts in hate speech classification, focusing on MBTI-based traits. A human annotation survey confirms that MBTI dimensions significantly affect labeling behavior. Extending this to LLMs, we prompt four open-source models with MBTI personas and evaluate their outputs across three hate speech datasets. Our analysis uncovers substantial persona-driven variation, including inconsistencies with ground truth, inter-persona disagreement, and logit-level biases. These findings highlight the need to carefully define persona prompts in LLM-based annotation workflows, with implications for fairness and alignment with human values.

仇恨言论大模型人格提示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。