arXiv:2606.09092cs.LG2026-06中稿 · ICML

用强化学习提升大模型心智理论能力,避免表面捷径陷阱

From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning

论文配图:From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning
图 1 · 摘自论文原文
  • 设计新框架检测心智理论数据集中的虚假关联捷径
  • 强化推理训练使心智理论能力提升6%-10%,尤其在复杂推理中表现突出
  • 适合关注大模型安全推理与真实理解能力的研究者

心智理论(ToM)是现代基础模型在真实世界中有效且安全运行的必备能力。近期工作通过后训练提升ToM,但我们发现其进展常被普遍存在的‘捷径’问题干扰:模型仅依赖虚假因果关联即可达到99%准确率,造成假象。为此,我们首先开发框架系统检测ToM数据集中的捷径,并指导未来数据构建。发现仅需状态追踪的问题(如“信念”)比需要深层推理的问题(如“意图”)更易出现捷径。基于四个无捷径数据集(覆盖三个ToM场景),我们全面评估结合可验证奖励与显式推理链的强化微调(Thinking-RFT)是否优于监督微调(SFT)。结果表明:Thinking-RFT在所有场景均有效提升ToM,平均比SFT高6%,复杂高阶推理中提升达10%,多模态任务中提升7%;对未见领域和高阶问题泛化更强,且对反事实情境更鲁棒。此外,推理与强化学习的协同作用至关重要,思维型强化微调(Thinking-RFT)比非思维型高出7%。强化学习通过锚定关键词、状态变化等因果线索来引导推理。本研究为构建有效、稳健的ToM后训练数据与能力提供重要参考。

原文摘要 · Abstract (English)

Theory of Mind (ToM) is a must-acquire skill for modern foundation model systems to operate effectively and safely in the real world. Recent works have explored honing ToM via post-training; however, we show that such progress is confounded by a pervasive "shortcut" issue: tasks can reach up to 99% accuracy by simply exploiting spurious causal correlations, leading to a false sense of ToM. Motivated by this, we first develop a framework to systematically examine ToM datasets for shortcuts and provide guidance for future development. We find that questions reducible to pure state tracking, such as "belief," are especially shortcut-prone compared to mind questions, such as "intention," where reasoning beyond tracking is required. Using four shortcut-free datasets across three ToM contexts, we then comprehensively study whether Reinforcement Fine-Tuning with verifiable rewards and explicit reasoning chains, called Thinking-RFT, elevates ToM beyond Supervised Fine-Tuning, or SFT. Our key findings are as follows. First, Thinking-RFT effectively improves ToM in all scenarios, with a 6% improvement over SFT, particularly in complex higher-order reasoning, with a 10% improvement over SFT, and multimodal cases, with a 7% improvement over SFT. It also generalizes notably better to unseen domains and higher-order queries while being more robust to counterfactuals. Second, ToM benefits specifically from the joint effect of reasoning and RL: Thinking-RFT outperforms Non-Thinking-RFT by 7% on average. Third, RFT works by learning to ground its reasoning on anchor cues, such as keywords and state changes, that correspond to causal factors. We believe our study is useful for developing effective and robust ToM post-training datasets and advancing critical ToM capabilities.

心智理论强化学习大模型推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。