用因子分析发现大模型能力有少数核心技能决定。
From Benchmarks to Skills: Low-Rank Factors for LLM Evaluation
- 通过因子分析揭示模型表现矩阵的低秩结构。
- 44个基准任务主要由少数共享技能解释,存在大量冗余。
- 可用来识别冗余任务、快速评估新模型或匹配特定技能需求。
当前大语言模型(LLMs)的评估高度依赖不断增长的基准测试集和综合得分,但这些分数究竟反映什么能力仍不明确。本文提出一种新评估范式,探究基准性能是否源于众多独立能力,还是仅依赖少数共享维度。我们对60个模型与44个基准的性能矩阵(60×44)应用因子分析(FA),发现其具有内在低秩结构——少数潜在因子即可捕捉大部分任务空间特征。该低秩几何表明现有任务间存在显著冗余,解释了为何多个基准看似测量相似能力。进一步发现,这些潜在因子对应于连贯的、类技能的行为维度。基于此潜在技能空间,我们开发三项实用工具:(i) 识别冗余任务,(ii) 用少量任务快速建模新模型能力,(iii) 选择符合特定技能需求的模型。本方法为单一综合得分提供了可靠替代方案,构建了一个可解释且实用的框架,用于理解与评估大模型的核心能力。
原文摘要 · Abstract (English)
Current evaluations of large language models (LLMs) rely heavily on a growing collection of benchmarks and on aggregate benchmark scores, yet it remains unclear what this comparison actually captures, and what these scores reveal about models' underlying capabilities. Here, we propose a new paradigm for LLM evaluation, by asking whether benchmark performance reflects many independent abilities, or rather relies on a small number of shared dimensions. To answer this, we apply Factor Analysis (FA) to a massive performance matrix of LLMs versus benchmarks \((60\times44)\) revealing an \emph{intrinsically low-rank} structure of that matrix. That is, a small number of latent factors captures most of the structure in the full task space. This low-rank geometry reveals substantial redundancy across existing tasks and explains why many benchmarks appear to be measuring overlapping abilities. We further show that these latent factors correspond to coherent, skill-like, dimensions of LLM behavior. Leveraging this latent skill-space, we deliver three practical tools for LLM evaluation and downstream users: (i)~identifying redundant tasks, (ii)~profiling new models using a small subset of tasks, and (iii)~selecting models aligned with desired skill profiles. Our method provides a solid alternative to the de-facto standard of a single aggregate score, and establishes an interpretable and practical framework for understanding and benchmarking LLM core capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。