arXiv:2607.11363cs.CLcs.AI2026-07被引 1

用对话博弈测试大模型的社交推理能力,发现多数模型在真实情境下仍表现不足。

Beyond Sally-Anne: Evaluating Theory of Mind in LLMs using Epistemic Schelling Points

论文配图:Beyond Sally-Anne: Evaluating Theory of Mind in LLMs using Epistemic Schelling Points
图 1 · 摘自论文原文
  • 设计双模型对话任务,通过语义聚焦点检验认知透明度下的协作能力
  • 仅顶尖模型能应对不同认知状态,多数失败源于对私有与共有知识混淆
  • 揭示传统测试无法发现的社交推理短板,适合关注模型社会智能的研究者

大型语言模型(LLMs)的理论心智(ToM)评估常依赖类萨莉-安妮任务,但这类测试易被预训练数据中的相似任务“破解”,且难以衡量模型在自然场景中的功能性ToM。为此,我们提出认知不对称谢林任务(EAST),一种双模型对话游戏,要求模型在不同认知透明度条件下独立达成语义聚焦点,以评估其鲁棒且可泛化的ToM能力。结果显示,只有前沿模型能有效应对变化的认知需求;分析推理轨迹表明,协作失败主要源于认知追踪错误,如混淆私人知识与共有知识。尽管在静态基准上表现良好,但稳健的社会推理和认知追踪仍是关键瓶颈,为未来模型评估与发展提供了明确方向。

原文摘要 · Abstract (English)

Text-based evaluations of Theory of Mind (ToM) in Large Language Models (LLMs) often involve cognitive tests akin to the Sally-Anne task that can be gamed due to exposure to relevantly similar tasks in pre-training and do not obviously test models' functional ToM abilities in ways that generalize to naturalistic settings. To address these issues, we introduce the Epistemic Asymmetry Schelling Task (EAST), a two-player dialogue game designed to benchmark robust and generalizable ToM abilities. By requiring LLM-LLM dyads to independently converge on semantic Schelling points under varying states of epistemic transparency, we evaluate whether models can robustly apply ToM to achieve coordination. Our results reveal a significant capability gap in functional social reasoning, with only frontier models successfully navigating the varying epistemic demands of the tasks. Analysis of reasoning traces shows that coordination failures are primarily driven by epistemic tracking errors, such as conflating private knowledge with mutual knowledge. Despite high performance on traditional static benchmarks, our study shows that robust social reasoning and epistemic tracking remain a critical bottleneck, providing concrete targets for future LLM evaluation and development.

理论心智社交推理语言模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。