为大模型评估构建分层能力体系,打破任务碎片化困局
From Isolated Tasks to Structured Capabilities: A Multilayer Taxonomy for Large Language Models

- 基于认知科学构建三层14大能力域91个子技能体系
- 发现语言理解与推理是研究热点,6个领域关注度不足2%
- 首次量化能力共现强度,揭示心智理论等高价值方向
大模型评估涵盖多种任务和基准,但现有研究仍以任务为中心组织,导致跨研究比较困难、能力覆盖盲区难以识别。本文提出一个多层能力分类体系,包含14个能力领域与91个子技能,分为基础、构建与整合三层。该体系基于人类认知科学定义能力,不依赖模型架构。层级划分依据发展顺序与功能支持假设,并将人类概念适配至可观测模型行为。为验证实用性,我们筛选了2023至2025年ACL、AAAI、ICML、NeurIPS共31,505篇论文,通过多模型标注、共识与仲裁,映射出15,934篇聚焦大模型的研究。研究注意力高度集中于语言语义能力(3,551篇,22.3%)、推理(3,388篇,21.3%)、规划与决策(2,149篇,13.5%)及感知(1,954篇,12.3%),而六项能力在论文中出现频率低于2%。各领域中最常见子技能的中位覆盖率达97.9%,且在10个领域中至少90%论文涉及。语言语义与推理组合出现频次最高(n=1,864,11.7%),相对提升2.47倍;心智理论与社会推理互动对的提升幅度最大(n=62,提升30.84倍)。通过将分析单位从孤立任务转向结构化能力,该分类体系助力研究组织、覆盖审计、评估解读与可检验的诊断、训练与迁移假说。
原文摘要 · Abstract (English)
Large language model (LLM) evaluation spans diverse tasks and benchmarks, yet evidence remains organized around tasks rather than the capabilities they probe. This fragmentation limits cross-study comparison, obscures capabilities tasks recruit, and makes coverage gaps difficult to identify. We introduce a multi-layer taxonomy of 14 capability domains and 91 subskills across Primitive, Constructed, and Integrative layers. Human cognitive science guides capability definition and organization, not LLM architecture. Layer assignments draw on developmental precedence and hypothesized functional support, while human-origin constructs are adapted to observable model behavior. To demonstrate operational utility, we screened 31,505 papers from ACL, AAAI, ICML, and NeurIPS between 2023 and 2025 and mapped 15,934 LLM-focused papers through multi-model annotation, consensus, and arbitration. Direct research attention concentrated on Language-Semantic Competence (3,551; 22.3%), Reasoning (3,388; 21.3%), Planning and Decision-Making (2,149; 13.5%), and Perception (1,954; 12.3%), whereas six domains appeared in fewer than 2% of papers. Within domains, the most frequent subskill had a median prevalence of 97.9% and appeared in at least 90% of papers in 10 of 14 domains. Language-Semantic Competence and Reasoning formed the highest-volume pair (n = 1,864; 11.7%; lift = 2.47), whereas Theory of Mind and Social Reasoning and Interaction showed the highest lift among pairs with at least 20 co-occurrences (n = 62; lift = 30.84). By shifting the unit of analysis from isolated tasks to structured capabilities, the taxonomy supports research organization, coverage audits, evaluation interpretation, and testable hypotheses for diagnosis, training, and transfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。