arXiv:2507.14355cs.CL2025-07被引 5

用真实对话评估大模型识别人格特质能力,发现效果有限。

Can LLMs Infer Personality from Real World Conversations?

  • 基于555份真实访谈构建评测基准,测试大模型预测人格
  • 最高相关性仅0.27,模型易偏向中高分预测,准确性不足
  • 思维链提示和长上下文略有改善,但整体仍难用于心理应用

大型语言模型(如GPT-4.1 Mini、Meta-LLaMA和DeepSeek)在从开放式语言中进行可扩展的人格评估方面展现出潜力。然而,人格推断仍具挑战性,早期研究多依赖缺乏心理测量有效性的合成数据或社交媒体文本。本文引入一个包含555份半结构化访谈的真实世界基准,每份访谈附有BFI-10自评量表得分,用于评估基于LLM的人格推断性能。测试了三种先进模型在零样本提示下对BFI-10条目预测,以及在零样本和思维链提示下对五大性格特质的推断。所有模型均表现出高重测信度,但构念效度有限:与真实得分的相关性较弱(最大Pearson相关系数r = 0.27),评分者间一致性低(Cohen's κ < 0.10),且预测倾向中高分水平。思维链提示和更长输入上下文虽小幅提升分布一致性,但未显著提高特质层级准确性。结果凸显当前大模型人格推断的局限性,强调心理应用需基于实证开发。

原文摘要 · Abstract (English)

Large Language Models (LLMs) such as OpenAI's GPT-4 and Meta's LLaMA offer a promising approach for scalable personality assessment from open-ended language. However, inferring personality traits remains challenging, and earlier work often relied on synthetic data or social media text lacking psychometric validity. We introduce a real-world benchmark of 555 semi-structured interviews with BFI-10 self-report scores for evaluating LLM-based personality inference. Three state-of-the-art LLMs (GPT-4.1 Mini, Meta-LLaMA, and DeepSeek) were tested using zero-shot prompting for BFI-10 item prediction and both zero-shot and chain-of-thought prompting for Big Five trait inference. All models showed high test-retest reliability, but construct validity was limited: correlations with ground-truth scores were weak (max Pearson's $r = 0.27$), interrater agreement was low (Cohen's $κ< 0.10$), and predictions were biased toward moderate or high trait levels. Chain-of-thought prompting and longer input context modestly improved distributional alignment, but not trait-level accuracy. These results underscore limitations in current LLM-based personality inference and highlight the need for evidence-based development for psychological applications.

人格分析大模型评测心理测量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。