测试大模型认知能力是否可划分为稳定五维结构
CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models

- 构建13类任务的程序化基准,用多方法验证认知维度
- 多数任务相关性为正,但仅一半方差可由单一主轴解释
- 提示工程仅产生微弱对角效应,不支持稳定维度标签
大语言模型的认知评分常被概括为各能力维度的组合,这些维度应跨任务一致、对匹配干预有选择性响应,并能泛化至定义它们的模型之外。我们提出CogArena,一个围绕多方法框架构建的13范式程序生成基准,用于判断认知任务得分在五个理论驱动分组中是否值得赋予维度标签。在55个开源模型上,几乎所有范式相关性为正,单一主轴解释了约一半方差。组内优势小且依赖评分敏感度,在不同模型家族间不确定。在独立冻结、全交叉研究中(12个模型,6个家族),针对性支架仅表现出微弱匹配组优势,无支架特异性对比通过多重检验校正,选择性也未提升跨家族预测效果。冻结确认标准失败。事后替代表述复现得到更小的正向估计,仍未能通过。综上,结果支持边界结论:理论对齐提示仅产生微弱的电池内对角倾向,现有证据无法确立稳定的五维认知结构。CogArena提供了一套工作流,在赋予模型得分以认知标签前,整合行为特征、协方差、匹配干预与跨家族预测。
原文摘要 · Abstract (English)
LLM cognitive scores are increasingly summarized as per-ability profiles whose dimensions should converge across tasks, respond selectively to matched interventions, and generalize beyond the models used to define them. We introduce CogArena, a procedurally generated 13-paradigm benchmark built around a multimethod framework for determining when cognitive-task scores warrant dimensional labels across five theory-motivated groupings. Across 55 open-weight models, nearly all paradigm correlations are positive and a common axis explains about half the variance. The within-grouping advantage is small, scoring-sensitive, and uncertain across model families. In a separately frozen, fully crossed study across 12 models from six families, targeted scaffolds show a small matched-grouping advantage, but no scaffold-specific contrast survives multiplicity correction and selectivity does not improve held-out-family prediction. The frozen confirmation criterion fails. A post-hoc alternate-wording replication produces a smaller positive estimate and again fails. Together, these results support a boundary conclusion. Theory-aligned prompting produces a small in-battery diagonal tendency, but the present evidence does not establish stable five-dimensional profiles. CogArena provides a workflow joining behavioral signatures, covariance, matched interventions, and out-of-family prediction before cognitive labels are attached to model scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。