arXiv:2507.15788cs.LGcs.AI2025-07

小模型通过强化学习学不会通用心智理论,反而会钻数据漏洞。

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning

  • 用强化学习训练小模型理解他人想法
  • 模型在训练集上表现好,但换任务就失效
  • 长期训练导致死记硬背数据模式,不适合真实社交场景

近期大语言模型在复杂推理上的突破主要得益于后训练阶段的规则化强化学习(RL)。这引发疑问:类似方法能否让小模型获得更细腻的人类社会智能,如心智理论(ToM)?本文系统评估了小规模模型通过可验证奖励强化学习(RLVR)获取通用ToM能力的可能性。我们在多个主流ToM数据集(HiToM、ExploreToM、FANToM)上训练模型,并在保留数据集(如OpenToM)上测试泛化能力。结果表明,小模型难以建立通用的ToM能力:虽然在分布内任务上性能提升,但无法迁移到特征不同的新任务。此外,长时间强化学习使模型“破解”训练数据的统计规律,导致分布内表现显著提升,而分布外任务性能无改善甚至下降。这说明其行为是窄域过拟合,而非真正抽象的心智理论掌握。

原文摘要 · Abstract (English)

Recent advancements in large language models (LLMs) have demonstrated emergent capabilities in complex reasoning, largely spurred by rule-based Reinforcement Learning (RL) techniques applied during the post-training. This has raised the question of whether similar methods can instill more nuanced, human-like social intelligence, such as a Theory of Mind (ToM), in LLMs. This paper investigates whether small-scale LLMs can acquire a robust and generalizable ToM capability through RL with verifiable rewards (RLVR). We conduct a systematic evaluation by training models on various combinations of prominent ToM datasets (HiToM, ExploreToM, FANToM) and testing for generalization on held-out datasets (e.g., OpenToM). Our findings indicate that small LLMs struggle to develop a generic ToM capability. While performance on in-distribution tasks improves, this capability fails to transfer to unseen ToM tasks with different characteristics. Furthermore, we demonstrate that prolonged RL training leads to models ``hacking'' the statistical patterns of the training datasets, resulting in significant performance gains on in-domain data but no change, or degradation of performance on out-of-distribution tasks. This suggests the learned behavior is a form of narrow overfitting rather than the acquisition of a true, abstract ToM capability.

心智理论强化学习小模型过拟合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。