arXiv:2510.18212cs.AIcs.LG2025-10被引 34

给通用人工智能下量化定义,用人类认知模型评估AI能力差距。

A Definition of AGI

  • 基于人类认知理论,将通用智能拆解为10个核心能力域。
  • 当前AI在知识类任务表现强,但长期记忆等基础能力严重不足。
  • 首次给出可量化的AGI得分,如GPT-4仅27%,适合评估研究者参考。

当前人工智能缺乏对通用智能(AGI)的明确界定,导致与人类认知水平之间的差距模糊不清。本文提出一个可量化的框架,将AGI定义为达到受过良好教育成人的认知多样性和熟练度。该方法以卡特尔-霍恩-卡罗尔(Cattell-Horn-Carroll)认知理论为基础,将通用智能分解为十个核心认知领域,包括推理、记忆和感知,并采用经验证的人类心理测量工具评估AI系统。应用该框架发现,现有模型的认知能力呈现高度“锯齿状”分布:在知识密集型任务中表现优异,但在基础认知机制上存在关键缺陷,尤其是长期记忆存储能力。由此得出的AGI得分(如GPT-4为27%,GPT-5为57%)清晰量化了进展速度与尚未跨越的巨大鸿沟。

原文摘要 · Abstract (English)

The lack of a concrete definition for Artificial General Intelligence (AGI) obscures the gap between today's specialized AI and human-level cognition. This paper introduces a quantifiable framework to address this, defining AGI as matching the cognitive versatility and proficiency of a well-educated adult. To operationalize this, we ground our methodology in Cattell-Horn-Carroll theory, the most empirically validated model of human cognition. The framework dissects general intelligence into ten core cognitive domains-including reasoning, memory, and perception-and adapts established human psychometric batteries to evaluate AI systems. Application of this framework reveals a highly "jagged" cognitive profile in contemporary models. While proficient in knowledge-intensive domains, current AI systems have critical deficits in foundational cognitive machinery, particularly long-term memory storage. The resulting AGI scores (e.g., GPT-4 at 27%, GPT-5 at 57%) concretely quantify both rapid progress and the substantial gap remaining before AGI.

通用智能认知评估量化指标AI评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。