arXiv:2605.02897cs.HCcs.AI2026-05

顶尖大模型人格趋于统一,更理性中立,少情绪化。

Same Voice, Different Lab: On the Homogenization of Frontier LLM Personalities

论文配图:Same Voice, Different Lab: On the Homogenization of Frontier LLM Personalities
图 1 · 摘自论文原文
  • 用144个性格维度评估前沿大模型,采用外部评分体系。
  • 所有模型均倾向理性、系统性表达,抑制愧疚、谄媚等特质。
  • 创意类模型也趋向中性,揭示行业对理想助手的共识。

LLM助手的人格在用户体验和响应质量感知中起关键作用。我们通过外部ELO基准性格评分体系,对144种性格特质进行了大规模实验,评估前沿大模型的人格表现。结果发现,所有测试模型均趋同于一种系统化、条理清晰且分析型的性格表达,同时压制如懊悔或奉承等特质。此外,模型在“分布中间特质”(如诗意、活泼)上的差异更大,但即便是这些所谓的“创造性”模型,其人格也趋向中性。这些趋同现象表明,一种隐性的最优助手行为标准正在自发形成。在训练方法多样的背景下,人格训练反而表现出高度一致性,揭示了模型开发者之间存在的默契共识。

原文摘要 · Abstract (English)

LLM assistant personalities play a critical role in user experience and perceived response quality. We present a large-scale experiment of frontier LLM personalities using external ELO-based traits scoring across 144 traits. We find that all models tested converge on a form of trait expression that is systematic, methodical, and analytical and suppress traits such as remorseful and sycophantic. Moreover, models tend to diverge more in their expression of ``middle-of-distribution traits`` such as poetic or playful, but even these so-called ``creative`` models tend to have more neutral identities. These similarities suggest an implicit emergence of a standard of optimal assistant behavior. In a landscape of varied training methods, character training, therefore, stands out for its uniformity, offering insight into a tacit consensus between model developers.

大模型人格行为一致评测方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。