arXiv:2602.16085cs.CLcs.AI2026-02ACL被引 5

41个开源大模型揭示语言统计如何影响错误信念推理

Language Statistics and False Belief Reasoning: Evidence from 41 Open-Weight LMs

  • 用41个开源模型测试虚假信念任务,分析语言统计对心智推理的影响
  • 34%的模型表现出对隐含知识状态的敏感性,更大模型表现更优
  • 发现人类与模型共享一种基于非事实动词的信念偏误,支持语言统计解释

语言模型(LMs)中的心理状态推理研究有助于理解人类社会认知理论——例如心智推理部分源于语言暴露——以及对语言模型自身能力的认知。然而,现有大量研究依赖少量封闭源代码模型,限制了对心理理论的严格检验和对模型能力的评估。本文通过评估来自不同模型家族的41个开源权重模型在虚假信念任务中的行为,复制并扩展了已有工作。结果显示,在34%的模型中观察到对隐含知识状态的敏感性;但与先前研究一致,无一模型能完全像人类一样“消除”该效应。更大的模型展现出更高的敏感性和更强的心理测量预测力。此外,我们利用模型行为提出并验证了一个关于人类认知的新假设:无论是人类还是模型,当知识状态由非事实动词(如‘约翰认为……’)提示时,都更倾向于归因于错误信念,而间接提示(如‘约翰看向……’)则较少引发此类偏误。与知识状态的主要效应相比,人类对提示语的敏感性虽高于模型,但其效应量分布却与模型效应量分布重叠——表明语言的分布统计可能解释人类的后一种现象,但无法解释前一种现象。这些结果展示了使用大规模开源模型测试人类认知理论和评估模型能力的价值。

原文摘要 · Abstract (English)

Research on mental state reasoning in language models (LMs) has the potential to inform theories of human social cognition--such as the theory that mental state reasoning emerges in part from language exposure--and our understanding of LMs themselves. Yet much published work on LMs relies on a relatively small sample of closed-source LMs, limiting our ability to rigorously test psychological theories and evaluate LM capacities. Here, we replicate and extend published work on the false belief task by assessing LM mental state reasoning behavior across 41 open-weight models (from distinct model families). We find sensitivity to implied knowledge states in 34% of the LMs tested; however, consistent with prior work, none fully ``explain away'' the effect in humans. Larger LMs show increased sensitivity and also exhibit higher psychometric predictive power. Finally, we use LM behavior to generate and test a novel hypothesis about human cognition: both humans and LMs show a bias towards attributing false beliefs when knowledge states are cued using a non-factive verb (``John thinks...'') than when cued indirectly (``John looks in the...''). Unlike the primary effect of knowledge states, where human sensitivity exceeds that of LMs, the magnitude of the human knowledge cue effect falls squarely within the distribution of LM effect sizes-suggesting that distributional statistics of language can in principle account for the latter but not the former in humans. These results demonstrate the value of using larger samples of open-weight LMs to test theories of human cognition and evaluate LM capacities.

语言模型心智推理虚假信念认知机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。