arXiv:2504.01698cs.CLcs.AI2025-04被引 8

大模型在心智理论任务中表现好,但未必真懂人类推理。

Do Theory of Mind Benchmarks Need Explicit Human-like Reasoning in Language Models?

  • 用强化学习和监督微调训练不同规模模型,测试其心智推理能力。
  • 70亿参数模型经强化学习后能生成可解释的推理链,30亿以下则出现推理崩溃。
  • 仅靠输出正确性微调也能达到高分,说明无需真正模拟人类心理。

心智理论(ToM)是理解他人心理状态的能力,对人类社交智能至关重要,也是先进人工智能的关键能力。近期大语言模型(LLMs)在多个ToM基准上表现优异,引发疑问:这些基准是否必须依赖显式的人类式推理?我们通过强化学习(RL)和监督微调(SFT)对0.5B至7B参数的模型进行实验,在多个ToM数据集上评估其表现。结果表明,强化学习对大模型(7B)有显著提升,能产生高质量、可解释且可迁移的信念追踪推理;但在小模型(≤3B)中导致“推理崩溃”,即以极简、低意义的响应实现高准确率与泛化能力。令人意外的是,仅通过监督微调即可在各基准上获得竞争性且可泛化的性能,常达到或超越强化学习模型的准确率,尽管未明确训练其生成结构化推理过程。这揭示了基准准确率与实际推理机制之间的关键差异:当前的ToM基准可能无需显式的人类式心理模拟即可破解。当模型规模有限或训练信号仅关注输出正确性时,它们可能利用针对数据结构的替代规则来高效应对任务。

原文摘要 · Abstract (English)

Theory of Mind (ToM), the ability to attribute mental states to others, is fundamental for human social intelligence and a critical capability for advanced Artificial Intelligence. Recent advancements in Large Language Models (LLMs) have shown promising performance on ToM benchmarks, raising the question: Do these benchmarks necessitate explicit human-like reasoning processes, or can models succeed through alternative strategies? We investigate this question empirically by applying Reinforcement Learning (RL) and Supervised Fine-Tuning (SFT) to LLMs of varying scales (0.5B to 7B parameters) and evaluating them across multiple ToM datasets. Our results reveal a scale-dependent impact of RL: while RL significantly improves accuracy and fosters high-quality, interpretable, and transferable belief-tracking reasoning in larger models (7B), it leads to "reasoning collapse" in smaller models ($\leq$3B), where high accuracy and generalization ability are achieved via drastically shortened, less meaningful responses. Surprisingly, further SFT achieves competitive and generalizable performance across these benchmarks, often matching or exceeding RL models in accuracy, despite not being explicitly trained to produce structured reasoning traces. These findings highlight a critical discrepancy between benchmark accuracy and the nature of learned reasoning. Our work suggests that current ToM benchmarks may be solvable without requiring the explicit, human-like simulation of mental states they were designed to probe. LLMs, particularly when scale is limited or training signals focus solely on output correctness, may leverage alternative rules effective for benchmark data structures.

心智理论大模型推理机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。