测试大模型理论心理能力,发现其在复杂任务中表现脆弱。
Probing the Robustness of Theory of Mind in Large Language Models
- 构建68个分10类复杂度的理论心理评测任务
- 所有主流模型在复杂任务上准确率普遍偏低
- 对环境状态变化的认知是模型薄弱环节,适合研究社交推理
随着ChatGPT等大型语言模型(LLM)的成功,学界开始宣称这些模型具备类似人类的社会推理能力,尤其是理论心理(ToM)。已有研究使用心理学风格的任务验证了这一能力,但后续研究发现,当任务稍作改动时,能力即消失。本文提出一个包含68个任务的新数据集,涵盖10个复杂度等级,用于探测大模型在不同情境下的理论心理表现。我们评估了四款顶尖开源大模型在本数据集及Kosinski(2023)数据集上的表现。结果显示,所有模型整体目标准确率均较低,表明其理论心理能力有限。简单任务上各模型表现相近;但在涉及代理对环境自动状态变化的认知任务中,即使明确提示,所有模型仍表现不佳。当任务通过替换介词改变物体间关系时,所有模型性能下降,尤其以混合专家模型影响最显著。本研究为提升和稳定大模型的理论心理能力提供了方向。
原文摘要 · Abstract (English)
With the success of ChatGPT and other similarly sized SotA LLMs, claims of emergent human like social reasoning capabilities, especially Theory of Mind (ToM), in these models have appeared in the scientific literature. On the one hand those ToM-capabilities have been successfully tested using tasks styled similar to those used in psychology (Kosinski, 2023). On the other hand, follow up studies showed that those capabilities vanished when the tasks were slightly altered (Ullman, 2023). In this work we introduce a novel dataset of 68 tasks for probing ToM in LLMs, including potentially challenging variations which are assigned to 10 complexity classes. This way it is providing novel insights into the challenges LLMs face with those task variations. We evaluate the ToM performance of four SotA open source LLMs on our dataset and the dataset introduced by (Kosinski, 2023). The overall low goal accuracy across all evaluated models indicates only a limited degree of ToM capabilities. The LLMs' performance on simple complexity class tasks from both datasets are similar. Whereas we find a consistent tendency in all tested LLMs to perform poorly on tasks that require the realization that an agent has knowledge of automatic state changes in its environment, even when those are spelled out to the model. For task complications that change the relationship between objects by replacing prepositions, we notice a performance drop in all models, with the strongest impact on the mixture-of-experts model. With our dataset of tasks grouped by complexity we offer directions for further research on how to stabilize and advance ToM capabilities in LLM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。