arXiv:2608.09638cs.AIcs.CL2026-08

用游戏机制测评大模型的精细心理推理能力,发现模型懂规则但不会表达、靠训练而非思考。

Avalon-ToM-Bench: Evaluating Fine-Grained Theory of Mind via Asymmetric Game Mechanics

  • 通过游戏中的信息不对称设计,拆解心理推理为四类细粒度任务
  • 模型懂规则但推理弱,隐藏状态中已存正确推断却无法生成
  • 训练出的推理策略比临时思考更关键,适合研究社会智能的团队

心智理论(ToM)对智能体交互至关重要,但现有评估或依赖静态场景简化心理状态推理,或在互动环境中缺乏诊断性。我们提出Avalon-ToM-Bench,一个基于《抵抗:阿瓦隆》不对称信息机制的细粒度基准。不评估完整游戏表现,而是通过人工设计的视角受限问题,将ToM分解为认知与动机推理交叉推理与行动的2×2分类。对28个大模型的测试揭示三点:1)推理而非知识是瓶颈。模型具备良好的游戏规则理解,但心理推理能力明显不足,说明失败源于社交推理缺失而非知识缺乏;2)表达而非表征是关键。线性探测和激活操控分析显示,模型常在隐状态中正确表示心理推断,但生成阶段未能表达——线性探针准确率达77-82%,而模型自身链式思考仅62-70%;3)策略而非思辨更有效。专门的推理训练带来显著提升,而测试时链式思考仅小幅增益(平均+11.0对比+1.1),表明稳健的ToM依赖于习得的推理策略,而非增加推理时间的思辨。

原文摘要 · Abstract (English)

Theory of Mind (ToM) is essential for agent interactions, yet existing evaluations either rely on static scenarios that oversimplify mental-state reasoning or interactive settings that provide limited diagnostic insight. We present Avalon-ToM-Bench, a fine-grained benchmark that operationalizes ToM through the asymmetric-information mechanics of The Resistance: Avalon. Rather than evaluating end-to-end gameplay, it decomposes ToM into a 2$\times$2 taxonomy -- epistemic versus motivational reasoning crossed with inference versus action -- using human-crafted, perspective-constrained queries. Benchmarking 28 LLMs reveals three insights: 1) Reasoning, not knowledge. Models show strong game-rule comprehension but markedly weaker ToM abilities, isolating failures to social reasoning rather than missing domain knowledge. 2) Expression, not representation. Mechanistic analyses via linear probing and activation steering show that models frequently represent correct mental-state inferences in their hidden states but fail to express them during generation -- linear probes recover 77-82% accuracy versus 62-70% from the models' own chain-of-thought. 3) Policy, not deliberation. Dedicated reasoning training yields substantial improvements whereas test-time chain-of-thought provides only marginal gains (+11.0 versus +1.1 points on average), suggesting that robust ToM depends on a learned reasoning policy rather than increased inference-time deliberation.

心智理论大模型评测推理机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。