arXiv:2502.12566cs.AI2025-02被引 22

给大模型赋予不同性格,能显著影响其输出的偏见与毒性。

Exploring the Impact of Personality Traits on LLM Bias and Toxicity

  • 基于人格六因素框架设计实验提示,测试模型输出
  • 调整性格特质可有效降低模型输出的偏见与毒性
  • 为可控文本生成提供低成本新思路,适合安全研究者

随着AI在人类生活中扮演的角色日益多样,赋予大语言模型(LLMs)不同人格成为研究热点。尽管人格化提升了交互体验和适应性,却引发对内容安全的担忧,尤其是模型输出中的偏见、情绪倾向与毒性问题。本研究探讨将不同人格特质赋予LLMs如何影响其输出的毒性与偏见。基于社会心理学中广泛接受的HEXACO人格框架,我们设计了实验性提示,测试三类LLMs在三个毒性和偏见基准上的表现。结果表明,所有三类模型均对HEXACO人格特质敏感,且输出的偏见、负面情绪和毒性存在一致变化。尤其发现,调节若干人格特质水平可有效降低模型的偏见与毒性,类似于人类中人格与攻击行为的相关性。研究强调,在进行模型人格化时,需额外关注内容安全性,而人格调节或可作为实现可控文本生成的简单、低成本方法。

原文摘要 · Abstract (English)

With the different roles that AI is expected to play in human life, imbuing large language models (LLMs) with different personalities has attracted increasing research interests. While the "personification" enhances human experiences of interactivity and adaptability of LLMs, it gives rise to critical concerns about content safety, particularly regarding bias, sentiment and toxicity of LLM generation. This study explores how assigning different personality traits to LLMs affects the toxicity and biases of their outputs. Leveraging the widely accepted HEXACO personality framework developed in social psychology, we design experimentally sound prompts to test three LLMs' performance on three toxic and bias benchmarks. The findings demonstrate the sensitivity of all three models to HEXACO personality traits and, more importantly, a consistent variation in the biases, negative sentiment and toxicity of their output. In particular, adjusting the levels of several personality traits can effectively reduce bias and toxicity in model performance, similar to humans' correlations between personality traits and toxic behaviors. The findings highlight the additional need to examine content safety besides the efficiency of training or fine-tuning methods for LLM personification. They also suggest a potential for the adjustment of personalities to be a simple and low-cost method to conduct controlled text generation.

大模型安全人格化偏见控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。