arXiv:2601.12410cs.AI2026-01

测试大模型能否像人一样理解他人认知,发现表现远不如人类。

Are LLMs Smarter Than Chimpanzees? An Evaluation on Perspective Taking and Knowledge State Estimation

  • 设计两任务:预测角色行为与识别不合理知识
  • 主流大模型在两项任务中表现接近随机
  • 研究呼吁重视模型对意图和认知的理解能力

认知人类学认为,人类智能的关键在于推断他人知识状态和意图,而黑猩猩不具备此能力。本文旨在评估大语言模型在估计他人知识状态及潜在行为方面的表现。我们设计了两个任务:(1) 判断模型能否基于角色自身的知识预测其下一步行动,而非使用视角外的信息;(2) 判断模型能否识别角色通过行为表现出其不应拥有的知识。结果表明,当前最先进的大模型在两项任务中均表现接近随机水平,显著低于人类表现。研究认为,未来大模型研究应更重视知识状态估计与意图理解能力的提升。

原文摘要 · Abstract (English)

Cognitive anthropology suggests that the distinction of human intelligence lies in the ability to infer other individuals' knowledge states and understand their intentions. In comparison, our closest animal relative, chimpanzees, lack the capacity to do so. With this paper, we aim to evaluate LLM performance in estimating other individuals' knowledge states and their potential actions. We design two tasks to test (1) if LLMs can predict story characters' next actions based on their own knowledge vs. improperly using information unavailable from their perspective, and (2) if LLMs can detect when story characters, through their actions, demonstrate knowledge they should not possess. Results reveal that most current state-of-the-art LLMs achieve near-random performance on both tasks, and are substantially inferior to humans. We argue future LLM research should place more weight on the abilities of knowledge estimation and intention understanding.

认知推理大模型评估知识状态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。