用经典小说角色测试大模型对背景上下文的理解能力
The Essence of Contextual Understanding in Theory of Mind: A Study on Question Answering with Story Characters
- 构建包含1035个问题的CharToM基准,基于经典小说角色
- 人类读过小说时答题正确率显著高于未读者
- 即使预训练见过故事,大模型仍远逊于人类
心智理论(ToM)是人类理解他人心理状态的核心能力,依赖对人物背景和生活经历的综合理解。现有评测基准多使用缺乏全局背景的短篇叙事,忽视了人物背景信息的重要性。本文通过引入基于经典小说角色的CharToM基准,包含1035个ToM问题,验证了全面理解个人背景在ToM中的关键作用。人类实验显示,读过原著的参与者表现远超未读者。对当前主流大模型(包括o1和DeepSeek-R1)的测试表明,尽管这些模型在预训练中接触过相关文本,其表现仍显著落后于人类,暴露出当前大模型在捕捉复杂上下文信息以进行心智推理方面的局限性。
原文摘要 · Abstract (English)
Theory-of-Mind (ToM) is a fundamental psychological capability that allows humans to understand and interpret the mental states of others. Humans infer others' thoughts by integrating causal cues and indirect clues from broad contextual information, often derived from past interactions. In other words, human ToM heavily relies on the understanding about the backgrounds and life stories of others. Unfortunately, this aspect is largely overlooked in existing benchmarks for evaluating machines' ToM capabilities, due to their usage of short narratives without global context, especially personal background of characters. In this paper, we verify the importance of comprehensive contextual understanding about personal backgrounds in ToM and assess the performance of LLMs in such complex scenarios. To achieve this, we introduce CharToM benchmark, comprising 1,035 ToM questions based on characters from classic novels. Our human study reveals a significant disparity in performance: the same group of educated participants performs dramatically better when they have read the novels compared to when they have not. In parallel, our experiments on state-of-the-art LLMs, including the very recent o1 and DeepSeek-R1 models, show that LLMs still perform notably worse than humans, despite that they have seen these stories during pre-training. This highlights the limitations of current LLMs in capturing the nuanced contextual information required for ToM reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。