用神经心理学测试评估大模型认知能力,发现其与人类有差距
A Neuropsychologically Grounded Evaluation of LLM Cognitive Abilities
- 基于三大神经心理测试构建新评测基准NeuroCognition
- 模型在图像和复杂任务上表现下降,简单策略更有效
- 该基准能揭示模型与人类认知的差异,适合改进模型
大型语言模型(LLMs)在10个基准上表现出统一的“通用能力因子”,这一发现通过156个模型的因子分析得到验证。然而,它们仍难以完成对人类而言简单的任务。原因在于当前基准侧重任务完成度,未能探测基础认知能力。为此,我们引入了基于三个改编神经心理学测试的NeuroCognition基准:瑞文渐进矩阵(抽象关系推理)、空间工作记忆(保持与系统搜索)以及威斯康星卡片分类测试(认知灵活性)。评估显示,尽管模型在文本任务上表现良好,但在图像任务及复杂情境下性能下降。同时,复杂推理并非普遍有益,而简单的人类策略可带来部分提升。此外,NeuroCognition与标准通用能力基准呈正相关,但仍测量出超越它们的独特认知能力。总体而言,NeuroCognition突出了当前LLMs与人类智能的契合点与缺失的核心适应性认知,具有作为可验证、可扩展的改进来源潜力。
原文摘要 · Abstract (English)
Large language models (LLMs) exhibit a unified "general factor" of capability across 10 benchmarks, a finding confirmed by our factor analysis of 156 models, yet they still struggle with simple, trivial tasks for humans. This is because current benchmarks focus on task completion, failing to probe the foundational cognitive abilities that highlight these behaviors. We address this by introducing the NeuroCognition benchmark, grounded in three adapted neuropsychological tests: Raven's Progressive Matrices (abstract relational reasoning), Spatial Working Memory (maintenance and systematic search), and the Wisconsin Card Sorting Test (cognitive flexibility). Our evaluation reveals that while models perform strongly on text, their performance degrades for images and with increased complexity. Furthermore, we observe that complex reasoning is not universally beneficial, whereas simple, human-like strategies yield partial gains. We also find that NeuroCognition correlates positively with standard general-capability benchmarks, while still measuring distinct cognitive abilities beyond them. Overall, NeuroCognition emphasizes where current LLMs align with human-like intelligence and where they lack core adaptive cognition, showing the potential to serve as a verifiable, scalable source for improving LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。