评测大模型从文本预测人格特质的能力,发现效果不稳且易有偏见。
Mind Reading or Misreading? LLMs on the Big Five Personality Test
- 用增强提示提升输出质量,但引入了倾向“有特质”的系统性偏差。
- 开放性与宜人性较易识别,外向性和神经质仍难准确预测。
- 零样本下无模型表现稳定,需谨慎设计提示与评估指标。
我们评估了大语言模型(LLMs)在二分类五因素人格模型(BIG5)下,从文本自动预测人格特质的能力。测试涵盖五种模型(含GPT-4及轻量开源模型),三个异构数据集(Essays、MyPersonality、Pandora),以及两种提示策略(最小提示与包含语言及心理线索的增强提示)。增强提示虽减少无效输出并改善类别平衡,但引入了倾向预测特质存在的系统性偏差。性能差异显著:开放性与宜人性相对易检测,而外向性和神经质仍具挑战。尽管部分开源模型在某些设置下接近GPT-4及以往基准,但在零样本二分类场景中无配置能实现一致可靠的预测。此外,整体指标如准确率与宏平均F1掩盖了显著的类别不平衡问题,每类召回率提供更清晰的诊断价值。研究表明,当前现成大模型尚不适合人格预测任务,提示设计、特质表述与评估指标的协调至关重要。
原文摘要 · Abstract (English)
We evaluate large language models (LLMs) for automatic personality prediction from text under the binary Five Factor Model (BIG5). Five models -- including GPT-4 and lightweight open-source alternatives -- are tested across three heterogeneous datasets (Essays, MyPersonality, Pandora) and two prompting strategies (minimal vs. enriched with linguistic and psychological cues). Enriched prompts reduce invalid outputs and improve class balance, but also introduce a systematic bias toward predicting trait presence. Performance varies substantially: Openness and Agreeableness are relatively easier to detect, while Extraversion and Neuroticism remain challenging. Although open-source models sometimes approach GPT-4 and prior benchmarks, no configuration yields consistently reliable predictions in zero-shot binary settings. Moreover, aggregate metrics such as accuracy and macro-F1 mask significant asymmetries, with per-class recall offering clearer diagnostic value. These findings show that current out-of-the-box LLMs are not yet suitable for APPT, and that careful coordination of prompt design, trait framing, and evaluation metrics is essential for interpretable results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。