arXiv:2604.09911q-bio.NCcs.AI2026-04被引 1

AI模型的通用智能表现其实隐藏着能力分化,越强的模型越依赖工具分工。

The Rise and Fall of $G$ in AGI

论文配图:The Rise and Fall of $G$ in AGI
图 1 · 摘自论文原文
  • 用主成分分析将大模型在多个基准上的表现看作认知测试,发现普遍存在正相关结构
  • 2023–2024年核心基准上第一主成分解释92%方差,2024年后降至64%因推理专用模型出现
  • 模型通过外挂工具实现“推理”能力,导致通用因子(G)下降,体现智能从统一到分工的转变

心理学中‘一般智力’描述的是能力间的正相关关系,而非能力数量。本文将心理测量学中的斯皮尔曼g因子(衡量正相关网络)与人工智能通用智能(AGI)在时间结构化基准上的表现关联起来。将大语言模型发布视为被试,以模型×基准×时间矩阵(39个模型,2019–2025年;14个基准)为基础,进行主成分分析。初步结果表明,在8个基准上所有28组两两相关均为正,形成强正相关结构。对基准相关谱随时间演化分析发现,5个基准核心集在2024年前由第一主成分解释90%方差,后降至77%;四基准电池中,第一主成分峰值达92%(2023–2024),2024年推理专用模型出现后降为64%。此变化与‘推理’能力向工具外化同步发生,暗示底层智能结构从统一(‘智狐’)转向专业化(‘智猬’)。从严格心理测量学角度,大模型表现出通用智能压制特定智能的现象。这颠覆了‘以简驭繁’的理想,形成架构越来越复杂、能力逐步提升的‘托勒密式演替’。

原文摘要 · Abstract (English)

In the psychological literature the term `general intelligence' describes correlations between abilities and not simply the number of abilities. This paper connects Spearman's $g$-factor from psychometrics, measuring a positive manifold, to the implicit ``$G$-factor'' in claims about artificial general intelligence (AGI) performance on temporally structured benchmarks. By treating LLM benchmark batteries as cognitive test batteries and model releases as subjects, principal component analysis is applied to a models $\times$ benchmarks $\times$ time matrix spanning 39 models (2019--2025) and 14 benchmarks. Preliminary results confirm a strong positive manifold in which all 28 pairwise correlations positive across 8 benchmarks. By analyzing the spectrum of the benchmark correlation through time, PC1 explains 90\% of variance on a 5-benchmark core battery ($n=19$)) reducing to 77\% by 2024. On a four benchmark battery, PC1 is found to peak at 92\% of the variance between 2023--2024 and reduce to 64\% with the arrival of reasoning-specialized models in 2024. This is coincident with a rotation in the G-factor as models outsource `reasoning' to tools. The analysis of partial correlation matrices through time provides evidence for the evolution of specialization beneath the positive manifold of general intelligence (AI-hedgehog) encompassing diverse high dimensional problem solving systems (AI-foxes). In strictly psychometric terms, AI models exhibit general intelligence suppressing specialized intelligences. LLMs invert the ideal of substituting complicated models with parsimonious mechanisms, a `Ptolemaic Succession' of theories, with architectures of increasing hierarchical complication and capability.

通用智能大模型智能演化心理测量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。