大模型不能模拟人类心理,用作心理学实验参与者不可靠。
Large Language Models Do Not Simulate Human Psychology
- 通过概念与实证分析,揭示大模型响应受措辞微调影响显著。
- 同一心理任务中,大模型与人类回答差异明显,即使使用专门微调模型。
- 不同大模型对新问题反应不一,说明其结果不可靠,需人工验证。
大型语言模型(如ChatGPT)在研究中应用日益广泛,涵盖从写作辅助到复杂数据标注的任务。近期有研究认为,这些模型可能具备模拟人类心理的能力,可替代人类参与心理学实验。本文对此提出警告。我们从概念层面质疑该假设,并通过实证证据加以支持:仅微调措辞即引发语义大幅变化,导致大模型与人类回答出现显著差异,即便使用专为心理响应优化的CENTAUR模型亦然。此外,不同大模型对新样本的回应差异巨大,进一步表明其不可靠性。结论是,大模型不具备模拟人类心理的能力,建议心理学研究者将其视为有用但根本不可靠的工具,任何新应用都必须通过人类对照验证。
原文摘要 · Abstract (English)
Large Language Models (LLMs),such as ChatGPT, are increasingly used in research, ranging from simple writing assistance to complex data annotation tasks. Recently, some research has suggested that LLMs may even be able to simulate human psychology and can, hence, replace human participants in psychological studies. We caution against this approach. We provide conceptual arguments against the hypothesis that LLMs simulate human psychology. We then present empiric evidence illustrating our arguments by demonstrating that slight changes to wording that correspond to large changes in meaning lead to notable discrepancies between LLMs' and human responses, even for the recent CENTAUR model that was specifically fine-tuned on psychological responses. Additionally, different LLMs show very different responses to novel items, further illustrating their lack of reliability. We conclude that LLMs do not simulate human psychology and recommend that psychological researchers should treat LLMs as useful but fundamentally unreliable tools that need to be validated against human responses for every new application.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。