用大模型模拟人类心理,发现其在整体表现上像人,细节上仍有差距。
Psychometric Comparability of LLM-Based Digital Twins
- 构建心理可比性框架,从概念表征到普遍规律全面评估模型
- 整体水平相关性强,但具体题目表现偏差明显,尤其在语言和决策中
- 适合用于群体趋势分析,不适合个体精细心理推断
大语言模型作为人类被试的数字孪生体,其心理测量可比性尚不明确。本文提出一个涵盖概念表征与普遍规律的构念效度框架,以人类金标准为基准进行评测。在多项研究中,数字孪生体在聚合层面表现出高准确率和人格轮廓高度相关,但在题目层面相关性减弱。在词语联想测试中,模型网络展现出类人的小世界结构和理论一致的社区划分,但词汇选择和局部结构存在差异。在决策与情境化任务中,模型未能复现启发式偏差,表现出规范理性、方差压缩和有限的时间敏感性。丰富特征与特质相关条件可提升大五人格预测及普遍规律一致性,但模型不变性仍受限,存在部分配置解与持续载荷差异。在英汉自由文本任务中,高维特征数字孪生体更接近构念级叙事内容,但语言与个体差异依然存在。研究揭示,数字孪生体仅在构念、任务与推断层级与人类数据一致时才具实际价值。
原文摘要 · Abstract (English)
Large language models (LLMs) act as digital twins for human respondents, yet their psychometric comparability remains uncertain. We propose a construct validity framework spanning construct representation and the nomothetic span, benchmarking models against human gold standards. Across studies, digital twins achieved high aggregate-level accuracy and profile correlations, but showed attenuated item-level correlations. In word association tests, LLM networks exhibited humanlike small-world structure and theory-consistent communities, yet diverged lexically and in local structure. In decision-making and contextualized tasks, they under-reproduced heuristic biases, demonstrating normative rationality, compressed variance, and limited temporal sensitivity. Feature-rich and trait relevant conditioning improved Big Five personality prediction and nomothetic-span alignment, but network invariance remained limited, with partial configural solutions and persistent loading differences. In cross-language free-text tasks in English and Chinese, feature-rich digital twins better approximated construct-level narrative content, but linguistic and idiographic differences persisted. These findings clarify that digital twins are most useful within validated boundaries, where the construct, task and level of inference align with evidence from human data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。