推理模型在心理理论任务中表现更稳定,未必是真正理解他人想法。
Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning
- 用强化学习训练的推理模型应对提示变化更稳定。
- 在多种测试中,模型对提示扰动的鲁棒性显著提升。
- 适合关注模型可靠性而非认知能力的读者。
大型语言模型(LLMs)在心理理论(ToM)测试中表现强劲,引发对其底层能力本质的争议。与此同时,通过可验证奖励进行强化学习训练的推理型LLM在多个基准测试中表现出显著提升。本文通过新颖的机器心理实验改编及已有基准测试结果,考察此类推理模型在ToM任务中的行为。结果显示,这些模型在提示变化和任务扰动下表现出更强的鲁棒性。分析表明,这种优势至少部分源于模型在提示或任务变化下仍能更稳定地得出正确答案。这支持了基于鲁棒性的解释,而非存在新的、专门的ToM能力。
原文摘要 · Abstract (English)
Large language models (LLMs) have recently shown strong performance on Theory of Mind (ToM) tests, prompting debate about the nature and validity of the underlying capabilities. At the same time, reasoning-oriented LLMs trained via reinforcement learning with verifiable rewards have demonstrated notable improvements across a range of benchmarks. In this work, we examine the behavior of such reasoning models in ToM tasks using novel adaptations of machine psychological experiments together with results from established benchmarks. We observe that reasoning models consistently exhibit increased robustness to prompt variations and task perturbations. Our analysis suggests these gains come at least partly from models being more robust at reaching the correct answer under prompt and task variation. We read this as evidence for a robustness-based account rather than for a new ToM-specific ability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。