arXiv:2512.09149cs.CLcs.AI2025-12

用心理测试评估大模型的人格模拟能力,发现模型表现随训练进步而提升。

MindShift: Analyzing Language Models' Reactions to Psychological Prompts

  • 基于MMPI心理量表设计人格化提示,测试模型角色适应性。
  • 不同模型家族在人格模拟上表现差异显著,反映其心理建模能力不一。
  • 公开基准数据与代码,适合研究模型社会行为与心理建模的学者使用。

大型语言模型(LLMs)具备吸收并反映用户指定人格特质和态度的潜力。本研究通过严谨的心理测量方法探索这一潜力,采用心理学文献中最广泛使用的明尼苏达多项人格问卷(MMPI)来考察语言模型的行为表现。为评估模型对心理提示的敏感性和偏见,我们构建了具有不同特质强度的系列人格化提示,从而衡量模型对角色设定的遵循程度。研究提出名为MindShift的基准,用于评估语言模型的心理适应能力。结果表明,模型在角色感知方面有持续改进,这归因于训练数据集和对齐技术的进步。同时,不同模型类型与家族在心理测评响应上存在显著差异,说明其模拟人类人格特质的能力各不相同。MindShift的提示与评估代码将公开共享。

原文摘要 · Abstract (English)

Large language models (LLMs) hold the potential to absorb and reflect personality traits and attitudes specified by users. In our study, we investigated this potential using robust psychometric measures. We adapted the most studied test in psychological literature, namely Minnesota Multiphasic Personality Inventory (MMPI) and examined LLMs' behavior to identify traits. To asses the sensitivity of LLMs' prompts and psychological biases we created personality-oriented prompts, crafting a detailed set of personas that vary in trait intensity. This enables us to measure how well LLMs follow these roles. Our study introduces MindShift, a benchmark for evaluating LLMs' psychological adaptability. The results highlight a consistent improvement in LLMs' role perception, attributed to advancements in training datasets and alignment techniques. Additionally, we observe significant differences in responses to psychometric assessments across different model types and families, suggesting variability in their ability to emulate human-like personality traits. MindShift prompts and code for LLM evaluation will be publicly available.

心理建模大模型评估人格模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。