用可控故事测试大模型能否理解他人心理,发现多数模型表现不佳。
Language Models Might Not Understand You: Evaluating Theory of Mind via Story Prompting
- 构建可编程故事生成框架,精准控制角色视角与情节
- 多数模型在心理建模任务上准确率低于世界建模任务
- 模型更擅长推理人而非物的心理状态,易依赖早期情节
我们提出StorySim,一个可编程框架,用于合成生成故事以评估大语言模型(LLMs)的理论心理(ToM)和世界建模(WM)能力。与以往可能受预训练数据污染或依赖其他LLM生成的故事基准不同,StorySim通过高度可控的剧情板(Storyboard)生成新颖、组合式的故事提示,能精确操控角色视角与事件发展。基于此框架,我们设计了一阶与二阶理论心理任务,以及控制心智状态追踪能力的世界建模任务。对一系列LLMs的实验表明,大多数模型在世界建模任务上的准确率高于理论心理任务;当推理对象为人时,模型表现更优。此外,框架揭示了模型存在启发式行为及对故事早期事件的过度依赖。所有数据生成与评估代码均开源。
原文摘要 · Abstract (English)
We introduce StorySim, a programmable framework for synthetically generating stories to evaluate the theory of mind (ToM) and world modeling (WM) capabilities of large language models (LLMs). Unlike prior benchmarks that may suffer from contamination in pretraining data, or rely on an LLM for generation, StorySim produces novel, compositional story prompts anchored by a highly controllable Storyboard, enabling precise manipulation of character perspectives and events. We use this framework to design first- and second-order ToM tasks alongside WM tasks that control for the ability to track and model mental states. Our experiments across a suite of LLMs show that most models achieve higher accuracy on WM tasks than on ToM tasks, and that models tend to reason more accurately when the subject of reasoning is a person rather than an inanimate object. Additionally, our framework enabled us to find evidence of heuristic behavior and an over-reliance on earlier events in the story. All code for generating data and evaluations is freely available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。