arXiv:2506.02461cs.CL2025-06ACL

测试大模型在五种语言中理解他人心理的能力,发现其表现因语言而异。

XToM: Exploring the Multilingual Theory of Mind for Large Language Models

  • 构建多语言心智理论评测基准XToM,覆盖五种语言
  • 大模型在多语言理解上表现好,但心理推理能力随语言波动
  • 揭示模型跨语言心智推理的局限性,适合关注多语言AI的学者

心智理论(ToM)是人类社会认知的核心能力,指推断他人心理状态的能力。现有大模型的ToM评估主要集中于英语,忽视了语言多样性对认知的影响。这一局限引发关键问题:大模型能否在多元语言背景下进行心智推理?为此,我们提出XToM,一个经过严格验证的多语言评测基准,涵盖五种语言,并包含多样、情境丰富的任务场景。利用XToM,我们系统评估了多个大模型(如DeepSeek R1),发现尽管模型在多语言理解方面表现优异,但其心智推理能力在不同语言间存在显著差异。结果揭示了大模型在跨语言心智推理方面与人类认知之间的明显差距。

原文摘要 · Abstract (English)

Theory of Mind (ToM), the ability to infer mental states in others, is pivotal for human social cognition. Existing evaluations of ToM in LLMs are largely limited to English, neglecting the linguistic diversity that shapes human cognition. This limitation raises a critical question: can LLMs exhibit Multilingual Theory of Mind, which is the capacity to reason about mental states across diverse linguistic contexts? To address this gap, we present XToM, a rigorously validated multilingual benchmark that evaluates ToM across five languages and incorporates diverse, contextually rich task scenarios. Using XToM, we systematically evaluate LLMs (e.g., DeepSeek R1), revealing a pronounced dissonance: while models excel in multilingual language understanding, their ToM performance varies across languages. Our findings expose limitations in LLMs' ability to replicate human-like mentalizing across linguistic contexts.

心智理论多语言大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。