大模型仅凭五大人格数据,就能精准推断其他心理特质关联模式。
From Five Dimensions to Many: Large Language Models as Precise and Interpretable Psychological Profilers
- 用大模型将五大人格得分转化为语言摘要,再推理生成其他心理量表结果。
- 生成的量表相关性与真人数据高度一致(R² > 0.89),超越语义相似性预测。
- 适合心理学建模、人格分析及理解大模型推理机制的研究者使用。
个体心理特征间普遍存在关联。我们研究了大语言模型(LLMs)能否仅通过少量量化输入,建模人类心理特质间的相关结构。将816名参与者的大五人格量表回答作为提示,让不同大模型扮演其角色,生成另外九个心理量表的响应。结果显示,大模型生成的数据在量表间相关性模式上与真实人类数据高度一致(R² > 0.89),零样本表现显著优于基于语义相似性的预测,接近直接在数据集上训练的机器学习算法。分析推理过程发现,大模型采用系统性两阶段策略:第一阶段将原始大五得分转化为自然语言人格摘要,实现信息选择与压缩;第二阶段基于摘要进行目标量表推理。模型识别出与训练算法相同的关键词人格因子,但未能区分因子内题项重要性。压缩后的摘要并非冗余,而是编码了协同信息——将其与原始得分结合后可提升预测一致性,表明其捕捉到了特质互动的高阶模式。研究证明,大模型能通过抽象与推理,仅凭少量数据精确预测个体心理特质,既为心理模拟提供强大工具,也揭示其涌现的推理能力。
原文摘要 · Abstract (English)
Psychological constructs within individuals are widely believed to be interconnected. We investigated whether and how Large Language Models (LLMs) can model the correlational structure of human psychological traits from minimal quantitative inputs. We prompted various LLMs with Big Five Personality Scale responses from 816 human individuals to role-play their responses on nine other psychological scales. LLMs demonstrated remarkable accuracy in capturing human psychological structure, with the inter-scale correlation patterns from LLM-generated responses strongly aligning with those from human data $(R^2 > 0.89)$. This zero-shot performance substantially exceeded predictions based on semantic similarity and approached the accuracy of machine learning algorithms trained directly on the dataset. Analysis of reasoning traces revealed that LLMs use a systematic two-stage process: First, they transform raw Big Five responses into natural language personality summaries through information selection and compression, analogous to generating sufficient statistics. Second, they generate target scale responses based on reasoning from these summaries. For information selection, LLMs identify the same key personality factors as trained algorithms, though they fail to differentiate item importance within factors. The resulting compressed summaries are not merely redundant representations but capture synergistic information--adding them to original scores enhances prediction alignment, suggesting they encode emergent, second-order patterns of trait interplay. Our findings demonstrate that LLMs can precisely predict individual participants' psychological traits from minimal data through a process of abstraction and reasoning, offering both a powerful tool for psychological simulation and valuable insights into their emergent reasoning capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。