构建首个覆盖46类认知任务的思维理论基准,评估大模型是否具备人类心智理解能力。
CogToM: A Comprehensive Theory of Mind Benchmark inspired by Human Cognition for Large Language Models
- 基于人类认知理论设计46种任务,涵盖8000多条双语实例
- 22个模型测试显示性能差异大,部分关键维度仍存在瓶颈
- 揭示大模型与人类认知结构可能存在的根本差异
大型语言模型(LLMs)是否真正具备类人思维理论(ToM)能力引发广泛关注。然而现有评测基准多局限于假信念等狭窄范式,难以全面反映人类认知机制。我们提出CogToM,一个系统性、理论驱动的基准,包含超过8000个双语样本,覆盖46种认知范式,并由49名人类标注者验证。对22个代表性模型(包括GPT-5.1和Qwen3-Max等前沿模型)的系统评估揭示显著性能差异,暴露出特定维度的持续瓶颈。基于人类认知模式的进一步分析表明,大模型与人类认知结构可能存在本质差异。CogToM为探究大模型认知边界的演进提供了可靠工具与新视角。
原文摘要 · Abstract (English)
Whether Large Language Models (LLMs) truly possess human-like Theory of Mind (ToM) capabilities has garnered increasing attention. However, existing benchmarks remain largely restricted to narrow paradigms like false belief tasks, failing to capture the full spectrum of human cognitive mechanisms. We introduce CogToM, a comprehensive, theoretically grounded benchmark comprising over 8000 bilingual instances across 46 paradigms, validated by 49 human annotator.A systematic evaluation of 22 representative models, including frontier models like GPT-5.1 and Qwen3-Max, reveals significant performance heterogeneities and highlights persistent bottlenecks in specific dimensions. Further analysis based on human cognitive patterns suggests potential divergences between LLM and human cognitive structures. CogToM offers a robust instrument and perspective for investigating the evolving cognitive boundaries of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。