arXiv:2608.26291cs.AIq-bio.NC2026-08

用博弈实验测大模型能否像人一样揣测他人意图,发现部分模型表现超人类。

Assessing mentalization in humans and large language models

论文配图:Assessing mentalization in humans and large language models
图 1 · 摘自论文原文
  • 设计经济博弈+计算建模,挖掘大模型的推理策略。
  • GPT-5能根据对手水平调整思考深度,性能超过人类。
  • 不同模型能力差异大,提示词可显著提升推理效果。

心智化——即推断他人信念与意图以指导自身行为——是人类社会互动的核心认知能力。尽管大语言模型(LLMs)在心理理论任务中表现出类人行为,但其是否能通过心智化实现适应性决策尚不明确。本文通过两种经济博弈结合认知计算建模,揭示了不同LLM的潜在心智化策略。测试了四个模型家族(DeepSeek、GPT-4.1、GPT-5、Gemini 2.0 Flash)共2,099个代理,在面对不同复杂度对手时的表现,并评估一种旨在激发策略推理的提示方法的效果。结果与251名人类参与者进行对比。两个博弈任务中,各模型均展现出明显的心智化行为与计算特征,且因模型厂商和规模而异。策略提示普遍提升了表现,但在不同任务中提升幅度不同。尤其值得注意的是,GPT-5代理能灵活调整递归推理深度以应对更复杂的对手,最终表现优于人类。研究证明不同大模型在心智化能力上存在差异,同时强调认知计算建模是衡量人类与机器智能的有力工具。

原文摘要 · Abstract (English)

Mentalization - the ability to infer others' beliefs and intentions to guide one's own choices - is a key cognitive function underlying human social interactions. Large language models (LLMs) demonstrate behaviour consistent with humans on theory-of-mind tasks, yet whether these models can guide adaptive behaviour through mentalization is unknown. Here we use two economic games with cognitive computational modeling to uncover the latent strategies underlying mentalization in LLMs. We tested individual LLM agents across four model families, DeepSeek, GPT-4.1, GPT-5 and Gemini 2.0 Flash (N = 2,099), against opponents of varying sophistication and examined whether a prompting strategy designed to elicit strategic reasoning improved performance. We benchmarked results against human participants (N = 251) as a comparative measure. Across both games, LLMs showed clear behavioural and computational signatures of mentalizing that differed markedly by model provider and size. Strategic prompting generally improved performance by inducing more sophisticated reasoning, yet the extent of the benefit differed across the two tasks. Last, GPT-5 agents flexibly adapted their recursive depth of reasoning to increasingly sophisticated opponents, demonstrating superior performance to human participants. Collectively, we demonstrate different capacities for mentalization across LLMs, and highlight cognitive computational modeling as a formal method for assessing comparative intelligence across humans and machines.

心智化大模型博弈实验认知建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。