arXiv:2604.12191cs.AI2026-04被引 1

用细粒度能力评估代替单一得分,更精准诊断大模型真实水平。

Beyond Scores: Diagnostic LLM Evaluation via Fine-Grained Abilities

论文配图:Beyond Scores: Diagnostic LLM Evaluation via Fine-Grained Abilities
图 1 · 摘自论文原文
  • 基于认知理论构建多维能力分类,数学达35维,跨学科扩展至物理、化学等
  • 通过多维项目反应模型预测未见题目表现,跨基准AUC达0.77~0.86
  • 适用于针对性训练、能力导向选型和智能评测设计,突破传统评分局限

当前大语言模型评估将多样任务性能汇总为单一分数,掩盖了细粒度能力差异,限制了针对性优化与任务适配选择。为此,我们提出一种认知诊断框架,可估计模型在多个细粒度维度上的能力水平。针对数学领域,构建了基于认知理论与领域知识的35维能力分类体系;框架采用多维项目反应理论与项目-能力关联矩阵,实现对细粒度能力的估计,并据此预测在未见题目(基准测试题)上的表现。在41个模型上评估显示,该方法具备强效准则效度,跨基准能力估计一致,对未见题目的预测准确率在单个基准内AUC为0.80~0.89,跨基准为0.77~0.86,显著优于基线。该框架可泛化至科学领域,在物理(27维)、化学(58维)和计算机科学(12维)中均保持一致诊断性能。本工作建立了细粒度能力评估的系统性框架,有望用于定向训练、能力导向选型及能力感知的评测设计。

原文摘要 · Abstract (English)

Current evaluations of large language models aggregate performance across diverse tasks into single scores. This obscures fine-grained ability variation, limiting targeted model improvement and ability-guided selection for specific tasks. Motivated by this gap, we propose a cognitive diagnostic framework that estimates model abilities across multiple fine-grained dimensions. For mathematics, we construct a 35-dimensional ability taxonomy grounded in cognitive theory and domain knowledge. The framework employs multidimensional Item Response Theory with an item-ability association matrix to estimate fine-grained ability levels, which in turn enable prediction of performance on unseen items (questions of benchmark). Evaluated on 41 models, our approach demonstrates strong criterion validity, consistent ability estimates across benchmarks, and accurate prediction of unseen items with AUC ranging from 0.80 to 0.89 within benchmarks and from 0.77 to 0.86 across benchmarks, substantially exceeding trivial baselines. The framework generalizes across scientific domains, producing consistent diagnostic performance in physics (27 dimensions), chemistry (58 dimensions), and computer science (12 dimensions). This work establishes a principled framework for fine-grained assessment of abilities, with potential applications in targeted training, ability-guided model selection, and ability-aware benchmark design.

大模型评估细粒度诊断能力建模多维评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。