arXiv:2605.06815cs.AIcs.CV2026-05

AI模型认知能力发展不均,语言推理远强于视觉推理

Uneven Evolution of Cognition Across Generations of Generative AI Models

论文配图:Uneven Evolution of Cognition Across Generations of Generative AI Models
图 1 · 摘自论文原文
  • 用心理学量表评估生成式AI的认知结构,对比人类水平
  • 语言理解与记忆接近人类98%分位,视觉推理低于1%分位
  • 揭示模型架构对语言符号的偏好,适合研究AGI局限性

追求通用人工智能需要超越特定任务表现的可靠方法来评估模型的认知能力。本文引入心理测量框架,评估生成式AI的认知特征,将其与人类标准对比并追踪跨代演变。使用改编自韦氏成人智力量表的任务对领先多模态模型进行初步评估,发现其认知架构极不均衡:言语理解与工作记忆接近天花板(>98百分位),而视觉推理接近地板(<1百分位)。为突破人类基准限制,我们开发了人工智能智商(AIQ)基准,应用于六代两个模型家族,揭示显著但非对称的性能提升。特别发现模态间存在明显分离:抽象数量推理在语言形式下发展远快于视觉对应形式,表明模型架构更偏爱基于语言的符号操作;尽管抽象视觉推理有所进步,但视觉感知组织基本停滞。这些结果表明,生成式模型的认知能力正不均衡演化,仅靠规模与优化可能无法克服根本架构局限,难以实现均衡的人类级通用智能。

原文摘要 · Abstract (English)

The pursuit of artificial general intelligence necessitates robust methods for evaluating the cognitive capabilities of models beyond narrow task performance. Here, we introduce a psychometric framework to assess the cognitive profiles of generative AI, comparing them to human norms and tracking their evolution across generations. Initial evaluation of leading multimodal models using tasks adapted from the Wechsler Adult Intelligence Scale revealed a profoundly uneven cognitive architecture: near-ceiling performance in verbal comprehension and working memory (>$98^{\text{th}}$ percentile) contrasted with near-floor performance in perceptual reasoning (<$1^{\text{st}}$ percentile). To track developmental trajectories beyond human-normed limits, we developed the Artificial Intelligence Quotient (AIQ) Benchmark and applied it to six generations and two model families, revealing significant but asymmetric performance gains. Notably, we uncovered a sharp dissociation between modalities; abstract quantitative reasoning matured far more rapidly when presented linguistically compared to a visually analogous format, indicating an architectural bias towards language-based symbolic manipulation. While abstract visual reasoning improved, visual-perceptual organization remained largely stagnant. Collectively, these findings demonstrate that the cognitive abilities of generative models are evolving unevenly, suggesting that scaling and optimization approaches to AGI development alone may be insufficient to overcome fundamental architectural limitations in achieving balanced, human-like general intelligence.

认知评估多模态语言优势架构局限

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。