arXiv:2410.19599econ.GNcs.AI2024-10被引 49

LLMs在人类行为模拟中表现不佳,研究者需谨慎使用。

Take Caution in Using LLMs as Human Surrogates: Scylla Ex Machina

  • 用11-20元请求游戏测试LLM推理深度
  • 多数先进模型无法复现人类行为分布
  • 输入语言、角色设定影响结果,不可预测

近期研究表明大语言模型(LLMs)可表现出类人推理,在经济实验、调查与政治话语中与人类行为一致。这促使研究者提议将LLMs作为社会科学研究中的人类替代或模拟工具。然而,LLMs与人类存在根本差异:其依赖概率模式,缺乏具身经验与生存目标等塑造人类认知的核心要素。本文通过11-20元请求游戏评估LLM的推理深度,发现几乎所有先进方法均未能在多个模型中复现人类行为分布。失败原因多样且不可预测,涉及输入语言、角色设定及安全机制等因素。研究警示:在研究人类行为或作为人类模拟时,应谨慎使用LLMs。

原文摘要 · Abstract (English)

Recent studies suggest large language models (LLMs) can exhibit human-like reasoning, aligning with human behavior in economic experiments, surveys, and political discourse. This has led many to propose that LLMs can be used as surrogates or simulations for humans in social science research. However, LLMs differ fundamentally from humans, relying on probabilistic patterns, absent the embodied experiences or survival objectives that shape human cognition. We assess the reasoning depth of LLMs using the 11-20 money request game. Nearly all advanced approaches fail to replicate human behavior distributions across many models. Causes of failure are diverse and unpredictable, relating to input language, roles, and safeguarding. These results advise caution when using LLMs to study human behavior or as surrogates or simulations.

大模型评测行为模拟社会科学研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。