大模型心理画像多是测量工具造成的假象,非模型真实特质。
Apparent Psychological Profiles of Large Language Models are Largely a Measurement Artifact

- 模型差异主要源于答题倾向偏差,而非真实性格特征。
- 81%-90%的模型差异由偏差导致,人类仅9%-16%。
- 换题目就能改画像,适合研究测评方法可靠性的学者。
用于人类的心理测量工具正被广泛用于为大语言模型(LLMs)生成稳定的心理画像,影响其可用性、安全评估及作为人类研究代理的使用。我们基于正式心理测量框架发现,这些画像主要是测量误差所致。对56个指令微调的LLMs与大规模人类参照样本施以涵盖自评和行为任务的多类人格与风险偏好量表,得到四项发现:第一,模型间差异并非源于量表所测特质,而是由方向性响应偏差驱动——即无论题目内容如何,倾向于朝量表一端或某一标签作答;方差分解显示,81%-90%的模型间差异归因于该偏差,而人类仅为9%-16%。第二,偏差随模型能力提升而下降,但未被消除。第三,由于响应由偏差主导,量表的表观信度几乎完全由其响应正交性决定,即量表项目中特质与偏差方向相反的比例。第四,模型呈现的心理画像随题目选择变化,可通过选题人为构造。结果表明,大模型的“心理画像”是测量工具的产物,而非模型本身属性。鉴于现有心理工具常缺乏对大模型的完全正交性且可能无效,我们呼吁开发以响应正交性为核心的专用评估体系。
原文摘要 · Abstract (English)
Psychological instruments designed for humans are increasingly used to assign large language models (LLMs) stable psychological profiles that affect their usability, safety assessment, and use as proxies for human participants in research. Using a formal psychometric framework, we show that these profiles are largely a measurement artifact. Administering a battery of personality and risk-preference instruments spanning self-reports and behavioral tasks to 56 instruction-tuned LLMs alongside large human reference samples, we report four findings. First, differences between models are driven not by the traits an instrument targets but by a directional response bias, a tendency to respond toward one end of the scale, or one labeled option, regardless of item content; a variance decomposition attributes 81-90% of between-model variation to this bias, against 9-16% in humans. Second, the bias declines with model capability but is not eliminated by it. Third, because bias rather than trait drives responding, an instrument's apparent reliability is almost entirely predicted by its response orthogonality, a term we coin for the proportion of items for which trait and bias point in opposite directions. Fourth, the profile a model appears to have shifts with the items used and can be manufactured through item selection. These results demonstrate that the apparent psychological profiles of LLMs are artifacts of the instrument used to measure them, not properties of the models themselves. As instruments borrowed from human psychology are rarely fully orthogonal and may inherently lack validity for LLMs, we call for dedicated assessments centered on response orthogonality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。