arXiv:2409.15324cs.AIcs.HC2024-09被引 4

用心理测量学方法研究大模型人格,发现所谓‘认知幻象’可能只是错觉。

Cognitive phantoms in LLMs through the lens of latent variables

  • 通过对比人类与三类大模型的潜在人格结构,检验问卷有效性
  • 发现人类量表在大模型中无法有效测量相似特质
  • 提醒研究者警惕误读模型行为,避免陷入认知幻象

大型语言模型(LLMs)日益应用于真实场景,亟需更深入理解其行为。其规模与复杂性使传统评估方法失效,促使研究借鉴心理学方法。近期研究通过心理量表测试大模型,报告其表现出类人特质,可能影响行为表现。然而该方法存在效度问题:预设这些特质存在于大模型中,并假设为人设计的工具可准确测量。典型流程未充分考虑此问题,仅比较和解释平均得分。本研究通过两种经验证的人格量表,比较人类与三类大模型的潜在结构。结果表明,为人类设计的量表在大模型中无法有效测量类似构念,这些构念甚至可能根本不存在,强调必须对大模型响应进行心理测量分析,以避免追逐认知幻象。

原文摘要 · Abstract (English)

Large language models (LLMs) increasingly reach real-world applications, necessitating a better understanding of their behaviour. Their size and complexity complicate traditional assessment methods, causing the emergence of alternative approaches inspired by the field of psychology. Recent studies administering psychometric questionnaires to LLMs report human-like traits in LLMs, potentially influencing LLM behaviour. However, this approach suffers from a validity problem: it presupposes that these traits exist in LLMs and that they are measurable with tools designed for humans. Typical procedures rarely acknowledge the validity problem in LLMs, comparing and interpreting average LLM scores. This study investigates this problem by comparing latent structures of personality between humans and three LLMs using two validated personality questionnaires. Findings suggest that questionnaires designed for humans do not validly measure similar constructs in LLMs, and that these constructs may not exist in LLMs at all, highlighting the need for psychometric analyses of LLM responses to avoid chasing cognitive phantoms. Keywords: large language models, psychometrics, machine behaviour, latent variable modeling, validity

大模型心理测量认知幻象潜变量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。