发现大模型可能具备整合心智理论与语用推理的社交世界模型。
On Emergent Social World Models -- Evidence for Functional Integration of Theory of Mind and Pragmatic Reasoning in Language Models
- 通过类认知神经科学方法定位语言模型中的社会认知机制。
- 在7类心智理论任务上表现关联,支持功能整合假说。
- 适合研究大模型社会智能与认知涌现的学者参考。
本文探究语言模型是否在一般心智理论(ToM)与语言特定语用推理之间复用共享的计算机制,以回应语言模型是否存在可跨任务复用的‘社交世界模型’这一核心问题。基于行为评估与因果机制实验,采用受认知神经科学启发的功能定位方法,在比以往研究更大规模的本地化数据集上分析了模型在7个子类心智理论能力(Beaudoin等, 2020)上的表现。严格的假设检验结果提供了支持功能整合假说的初步证据,表明语言模型可能发展出相互关联的‘社交世界模型’,而非孤立的能力模块。本研究贡献了新的心智理论本地化数据、功能定位技术的改进方法,以及对人工系统中社会认知涌现的实证洞察。
原文摘要 · Abstract (English)
This paper investigates whether LMs recruit shared computational mechanisms for general Theory of Mind (ToM) and language-specific pragmatic reasoning in order to contribute to the general question of whether LMs may be said to have emergent "social world models", i.e., representations of mental states that are repurposed across tasks (the functional integration hypothesis). Using behavioral evaluations and causal-mechanistic experiments via functional localization methods inspired by cognitive neuroscience, we analyze LMs' performance across seven subcategories of ToM abilities (Beaudoin et al., 2020) on a substantially larger localizer dataset than used in prior like-minded work. Results from stringent hypothesis-driven statistical testing offer suggestive evidence for the functional integration hypothesis, indicating that LMs may develop interconnected "social world models" rather than isolated competencies. This work contributes novel ToM localizer data, methodological refinements to functional localization techniques, and empirical insights into the emergence of social cognition in artificial systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。