arXiv:2506.23055cs.LGcs.AI2025-06

用心理量表测试大模型是否理解人类心理概念。

Measuring How LLMs Internalize Human Psychological Concepts: A preliminary analysis

  • 设计43个心理量表,通过相似性分析评估模型对概念的重构能力。
  • GPT-4分类准确率达66.2%,显著高于GPT-3.5和BERT。
  • 模型语义相似度与人类答题相关性可验证,适合心理学与AI交叉研究者。

大型语言模型(如ChatGPT)在生成类人文本方面表现出色,但其对塑造人类思维与行为的心理概念的内化程度尚不明确。本文构建了一套量化框架,利用43个经过验证的心理学量表,评估大模型与人类心理维度之间的概念一致性。方法基于成对相似性分析,判断模型能否准确重构和分类问卷项目,并通过层次聚类比较结果簇结构与原始标签的一致性。结果显示,GPT-4分类准确率为66.2%,显著优于GPT-3.5的55.9%和BERT的48.1%,均远超随机基线31.9%。此外,GPT-4估计的语义相似度与多个心理量表中人类响应的皮尔逊相关系数具有显著关联。该框架为评估人-模型概念对齐提供新路径,揭示现代大模型可测量地逼近人类心理构念,有助于构建更可解释的AI系统。

原文摘要 · Abstract (English)

Large Language Models (LLMs) such as ChatGPT have shown remarkable abilities in producing human-like text. However, it is unclear how accurately these models internalize concepts that shape human thought and behavior. Here, we developed a quantitative framework to assess concept alignment between LLMs and human psychological dimensions using 43 standardized psychological questionnaires, selected for their established validity in measuring distinct psychological constructs. Our method evaluates how accurately language models reconstruct and classify questionnaire items through pairwise similarity analysis. We compared resulting cluster structures with the original categorical labels using hierarchical clustering. A GPT-4 model achieved superior classification accuracy (66.2\%), significantly outperforming GPT-3.5 (55.9\%) and BERT (48.1\%), all exceeding random baseline performance (31.9\%). We also demonstrated that the estimated semantic similarity from GPT-4 is associated with Pearson's correlation coefficients of human responses in multiple psychological questionnaires. This framework provides a novel approach to evaluate the alignment of the human-LLM concept and identify potential representational biases. Our findings demonstrate that modern LLMs can approximate human psychological constructs with measurable accuracy, offering insights for developing more interpretable AI systems.

心理建模大模型评估概念对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。