arXiv:2410.07391cs.AI2024-10被引 9

用人类智商测试对比大模型,发现其在语言理解上超人类,但看图推理能力极差。

The Cognitive Capabilities of Generative AI: A Comparative Analysis with Human Benchmarks

  • 用韦氏成人智力量表测试大模型认知能力,分语言理解、工作记忆和视觉推理三域。
  • 工作记忆和语言理解得分超98%人类,但视觉推理仅达0.1-10%人类水平。
  • 模型越新越大越强,但多模态模型仍无法有效理解图像信息。

本文将领先的大型语言模型与视觉语言模型在韦氏成人智力量表(WAIS-IV)上与人类表现进行对比,重点评估语言理解(VCI)、工作记忆(WMI)和知觉推理(PRI)三个认知领域。大多数模型在存储、检索和操作符号(如字母数字序列)方面表现出色,其工作记忆指数(WMI)达到或超过人类群体常模的99.5百分位。语言理解指数(VCI)也稳定在98百分位以上,表明其对词汇意义及语义关系的理解能力优异。然而,多模态模型在知觉推理指数(PRI)上表现极差,得分范围仅为0.1至10百分位,暴露出其在解释和推理视觉信息方面的根本缺陷。较小或较旧版本的模型表现更弱,说明训练数据规模、参数量及调优技术的进步显著提升了模型的认知能力。

原文摘要 · Abstract (English)

There is increasing interest in tracking the capabilities of general intelligence foundation models. This study benchmarks leading large language models and vision language models against human performance on the Wechsler Adult Intelligence Scale (WAIS-IV), a comprehensive, population-normed assessment of underlying human cognition and intellectual abilities, with a focus on the domains of VerbalComprehension (VCI), Working Memory (WMI), and Perceptual Reasoning (PRI). Most models demonstrated exceptional capabilities in the storage, retrieval, and manipulation of tokens such as arbitrary sequences of letters and numbers, with performance on the Working Memory Index (WMI) greater or equal to the 99.5th percentile when compared to human population normative ability. Performance on the Verbal Comprehension Index (VCI) which measures retrieval of acquired information, and linguistic understanding about the meaning of words and their relationships to each other, also demonstrated consistent performance at or above the 98th percentile. Despite these broad strengths, we observed consistently poor performance on the Perceptual Reasoning Index (PRI; range 0.1-10th percentile) from multimodal models indicating profound inability to interpret and reason on visual information. Smaller and older model versions consistently performed worse, indicating that training data, parameter count and advances in tuning are resulting in significant advances in cognitive ability.

认知评估大模型多模态智商测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。