arXiv:2602.10625cs.AIcs.CL2026-02被引 1

大模型在心理理论任务中,推理越慢越容易出错,依赖选项匹配而非真正理解。

To Think or Not To Think, That is The Question for Large Reasoning Models in Theory of Mind Tasks

  • 通过控制推理长度和动态调整,发现适度推理更有效
  • 模型在长回答时准确率显著下降,大预算反而损害表现
  • 适合研究社会认知、人机交互的学者关注其局限性

心智理论(ToM)评估模型是否能推断隐藏的心理状态,如信念、欲望和意图,这对自然社交互动至关重要。尽管大型推理模型(LRMs)在数学和编程的逐步推理上取得进展,但其是否能迁移到社会认知能力仍不明确。我们对九个先进语言模型进行了系统研究,对比推理型与非推理型模型在三个典型ToM基准上的表现。结果显示,推理模型并未持续优于非推理模型,有时表现更差。细粒度分析揭示三点:第一,慢思考会崩溃——响应越长,准确率越低,更大的推理预算反而降低性能;第二,适度且自适应的推理有益——限制推理长度可缓解失败,不同成功模式表明动态适应的必要性;第三,选项匹配捷径:当多项选择项被移除后,推理模型表现明显提升,表明其依赖选项匹配而非真实推理。我们设计了两种干预方法:慢到快(S2F)自适应推理和思想到匹配(T2M)捷径防范,进一步验证并缓解问题。所有结果表明,LRMs在形式推理(如数学、代码)上的进步无法完全迁移到ToM这一典型社会推理任务。结论是,实现稳健的ToM需发展超越现有推理方法的独特能力。

原文摘要 · Abstract (English)

Theory of Mind (ToM) assesses whether models can infer hidden mental states such as beliefs, desires, and intentions, which is essential for natural social interaction. Although recent progress in Large Reasoning Models (LRMs) has boosted step-by-step inference in mathematics and coding, it is still underexplored whether this benefit transfers to socio-cognitive skills. We present a systematic study of nine advanced Large Language Models (LLMs), comparing reasoning models with non-reasoning models on three representative ToM benchmarks. The results show that reasoning models do not consistently outperform non-reasoning models and sometimes perform worse. A fine-grained analysis reveals three insights. First, slow thinking collapses: accuracy significantly drops as responses grow longer, and larger reasoning budgets hurt performance. Second, moderate and adaptive reasoning benefits performance: constraining reasoning length mitigates failure, while distinct success patterns demonstrate the necessity of dynamic adaptation. Third, option matching shortcut: when multiple choice options are removed, reasoning models improve markedly, indicating reliance on option matching rather than genuine deduction. We also design two intervention approaches: Slow-to-Fast (S2F) adaptive reasoning and Think-to-Match (T2M) shortcut prevention to further verify and mitigate the problems. With all results, our study highlights the advancement of LRMs in formal reasoning (e.g., math, code) cannot be fully transferred to ToM, a typical task in social reasoning. We conclude that achieving robust ToM requires developing unique capabilities beyond existing reasoning methods.

心智理论推理模型社会认知语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。