评测大模型心理理论能力,要求显式建模角色信念结构。
OmniToM: Benchmarking Theory of Mind in LLMs via Explicit Belief Modeling

- 通过显式构建角色信念命题来评估模型心理理论能力
- 零样本下发现模型在知识获取与信念表征上存在瓶颈
- 适合研究大模型社会认知与心智推理的学者使用
心理理论(ToM)是推断他人知识、意图和情绪的能力,现有大语言模型(LLMs)评估多采用终点问答方式,仅依据最终答案评分,难以判断模型是否真正构建了支撑推理的心理状态表征,尤其在涉及分歧、演化或错误信念的情境中。为此,我们提出OmniToM基准,直接通过要求模型为叙事中所有相关角色显式建模信念结构来评估其能力。这些信念结构由信念命题组成——即关于世界或他人心理状态的最小陈述,使知识、意图、情绪与错误信念可在统一格式下分析。评估分两阶段:第一阶段“信念提取”,从故事中提取影响社交动态的相关信念;第二阶段“信念标注”,为每个信念分配七维标签,涵盖递归层级、真值状态、知识可及性、显式程度、内容类型、心理来源与上下文。数据源自895个来自ToMBench的故事语料,并新增22,343条标注信念命题,采用人类校准的LLM辅助标注流程。零样本测试显示,当前大模型存在角色特异性信念追踪瓶颈:在将叙事事实转化为角色信念与共享心理状态时,知识可及性与表征决策能力显著不足。
原文摘要 · Abstract (English)
Theory of Mind (ToM), the ability to infer others' knowledge, intentions, and emotions, is commonly evaluated in large language models (LLMs) using end-point question answering, where performance is judged solely by the final answer to a social reasoning query. This paradigm obscures whether the model actually constructs the underlying mental-state representations required for robust reasoning, particularly in scenarios involving divergent, evolving, or mistaken beliefs. In order to address this research gap, we introduce OmniToM, a benchmark that directly evaluates these representations by requiring explicit modeling of belief structures for all relevant actors within a narrative. These structures are composed of belief propositions: minimal statements of what an actor takes to be true about the world or another actor's mental state, allowing knowledge, intentions, emotions, and false beliefs to be analyzed in a common format. Models are evaluated in two stages: Stage 1: Belief Extraction, which extracts from the story the beliefs relevant to its social dynamics, and Stage 2: Belief Labeling, which assigns each belief a seven-dimensional schema label covering recursive order, truth status, knowledge access, explicitness, content type, mental source, and context. Built from 895 stories from the existing ToMBench story corpus and augmented with 22,343 labeled belief propositions, OmniToM uses a human-calibrated LLM-assisted annotation pipeline. Across diverse models in zero-shot evaluation, OmniToM reveals an actor-specific belief-tracking bottleneck: current LLMs struggle with the knowledge-access and representational decisions required to transform narrative facts into actors' beliefs and shared mental states.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。