推理增强模型在心理理论任务中更抗干扰,表现更稳定。
Reasoning Promotes Robustness in Theory of Mind Tasks
- 用可验证奖励训练的推理模型提升对提示变化的鲁棒性。
- 在心理理论任务中,模型对扰动的适应能力显著增强。
- 适合关注大模型社会认知评估可靠性的研究者阅读。
大型语言模型(LLMs)在心理理论(ToM)测试中表现出色,引发对其能力本质的讨论。同时,通过可验证奖励强化学习(RLVR)训练的推理型模型在多个基准上取得显著进步。本文研究此类推理模型在心理理论任务中的表现,采用新型机器心理学实验和既有基准数据集。结果表明,推理模型对提示变化和任务扰动具有更强的鲁棒性。分析显示,性能提升主要源于更稳定的解题能力,而非产生新的心理理论推理形式。这一发现对评估大模型的社会认知行为具有重要启示。
原文摘要 · Abstract (English)
Large language models (LLMs) have recently shown strong performance on Theory of Mind (ToM) tests, prompting debate about the nature and true performance of the underlying capabilities. At the same time, reasoning-oriented LLMs trained via reinforcement learning with verifiable rewards (RLVR) have achieved notable improvements across a range of benchmarks. This paper examines the behavior of such reasoning models in ToM tasks, using novel adaptations of machine psychological experiments and results from established benchmarks. We observe that reasoning models consistently exhibit increased robustness to prompt variations and task perturbations. Our analysis indicates that the observed gains are more plausibly attributed to increased robustness in finding the correct solution, rather than to fundamentally new forms of ToM reasoning. We discuss the implications of this interpretation for evaluating social-cognitive behavior in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。