arXiv:2603.04161cs.CL2026-03被引 1

发现大模型社会认知能力受思维状态词汇影响,且训练方式决定其表现。

Traces of Social Competence in Large Language Models

  • 通过贝叶斯逻辑回归测试17个开源模型在192种假信念任务中的表现。
  • 模型规模提升有助于表现,但关键在解释性训练会强化特定响应模式。
  • 定位到'think'词向量为影响社会推理的核心因素,适合研究模型心智理论。

假信念测试(FBT)是评估心智理论(ToM)及社会认知能力的主要方法。对于大语言模型(LLMs),由于数据污染、模型细节不足和控制不一致等问题,该测试的可靠性与解释力受限。本文通过在192种平衡的FBT变体(Trott et al., 2023)上测试17个开源模型,并采用贝叶斯逻辑回归分析模型规模与后训练对社会认知能力的影响。结果表明,模型规模提升虽有益,但非线性;揭示出解释命题态度(X thinks)会根本性改变回答模式。指令微调部分缓解此效应,而以推理为导向的进一步微调则加剧之。对OLMo 2训练全过程的案例分析显示,该交叉效应在预训练阶段即已出现,说明模型习得了与心理状态词汇相关的刻板反应模式,可能压倒其他情境语义。最后,向量操控实验确认‘think’向量是驱动观察到的FBT行为的因果因子。

原文摘要 · Abstract (English)

The False Belief Test (FBT) has been the main method for assessing Theory of Mind (ToM) and related socio-cognitive competencies. For Large Language Models (LLMs), the reliability and explanatory potential of this test have remained limited due to issues like data contamination, insufficient model details, and inconsistent controls. We address these issues by testing 17 open-weight models on a balanced set of 192 FBT variants (Trott et al., 2023) using Bayesian Logistic regression to identify how model size and post-training affect socio-cognitive competence. We find that scaling model size benefits performance, but not strictly. A cross-over effect reveals that explicating propositional attitudes (X thinks) fundamentally alters response patterns. Instruction tuning partially mitigates this effect, but further reasoning-oriented fine-tuning amplifies it. In a case study analysing social reasoning ability throughout OLMo 2 training, we show that this cross-over effect emerges during pre-training, suggesting that models acquire stereotypical response patterns tied to mental-state vocabulary that can outweigh other scenario semantics. Finally, vector steering allows us to isolate a think vector as the causal driver of observed FBT behaviour.

心智理论社会认知大模型假信念测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。