测试人格提示对大模型社会推理解释质量的影响
Persona Prompting as a Lens on LLM Social Reasoning
- 用模拟人设引导模型生成解释,观察其变化
- 人格提示提升分类准确率但降低解释质量
- 模型始终存在偏见,人设难以真正改变输出
在仇恨言论检测等社会敏感任务中,大语言模型(LLMs)生成解释的质量对用户信任和模型对齐至关重要。尽管人格提示(Persona Prompting, PP)被广泛用于引导模型生成符合用户特征的内容,但其对模型推理过程的影响仍不明确。本文通过包含词级解释标注的数据集,评估不同模拟人口统计学人设下模型解释与各群体人类标注的一致性,并分析PP对模型偏见和人类对齐的影响。在三个LLM上的评估发现:(1)PP在最主观的任务(仇恨言论检测)中提升了分类性能,但降低了解释质量;(2)模拟人设未能有效匹配真实人群,且跨人设间高度一致,表明模型对人设引导具有强抵抗性;(3)模型普遍存在一致的群体偏见,且倾向于过度标记内容为有害,无论是否使用PP。研究揭示了关键权衡:虽能提升分类表现,但人格提示常以牺牲解释质量为代价,且无法缓解深层偏见,提示应用需谨慎。
原文摘要 · Abstract (English)
For socially sensitive tasks like hate speech detection, the quality of explanations from Large Language Models (LLMs) is crucial for factors like user trust and model alignment. While Persona prompting (PP) is increasingly used as a way to steer model towards user-specific generation, its effect on model rationales remains underexplored. We investigate how LLM-generated rationales vary when conditioned on different simulated demographic personas. Using datasets annotated with word-level rationales, we measure agreement with human annotations from different demographic groups, and assess the impact of PP on model bias and human alignment. Our evaluation across three LLMs results reveals three key findings: (1) PP improving classification on the most subjective task (hate speech) but degrading rationale quality. (2) Simulated personas fail to align with their real-world demographic counterparts, and high inter-persona agreement shows models are resistant to significant steering. (3) Models exhibit consistent demographic biases and a strong tendency to over-flag content as harmful, regardless of PP. Our findings reveal a critical trade-off: while PP can improve classification in socially-sensitive tasks, it often comes at the cost of rationale quality and fails to mitigate underlying biases, urging caution in its application.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。