提出三维评估框架,系统衡量大模型的认知、领域与任务能力。
CDT: A Comprehensive Capability Framework for Large Language Models Across Cognition, Domain, and Task
- 基于认知理论构建三维度评估框架,覆盖认知、领域与任务
- 在数据集评估与选择中提升性能,通用与特定任务分别增1.6和2.2分
- 适合需要全面评估模型能力的研究者与开发者
大型语言模型(LLM)的快速发展显著提升了其能力,亟需超越任务特定基准的综合评估框架。然而,现有基准多聚焦单一能力,缺乏整体评估体系。为此,我们提出认知-领域-任务(CDT)框架,从三个维度全面衡量模型能力。在认知层面,引入卡特尔-霍恩-卡罗尔认知理论,细化模型能力分类。我们将CDT应用于数据集能力评估与数据选择两个方向。实验表明,我们的能力指标与下游性能高度相关,能有效支持数据集分析与构建。在数据选择实验中,通用与特定任务基准分别取得44.3和45.4分,较基线提升1.6和2.2分。结果验证了CDT的有效性与实用性。源代码与模型见https://github.com/Alessa-mo/CDT。
原文摘要 · Abstract (English)
Recent advances in Large Language Models (LLMs) have significantly enhanced their capabilities, highlighting the need for comprehensive evaluation frameworks that extend beyond task-specific benchmarks. However, existing benchmarks often focus on isolated abilities, lacking a holistic framework for assessing LLM capabilities. To address this gap, we propose the Cognition-Domain-Task (CDT) framework, which comprehensively measures a model's capabilities across three dimensions. We expand the scope of model capability definitions at the cognitive level by incorporating the Cattell-Horn-Carroll cognitive theory, refining the categorization of model capabilities. We apply CDT in two directions: dataset capability evaluation and data selection. Experiments show that our capability metrics correlate well with downstream performance and can support effective dataset analysis and construction. The experiments on data selection also show significant improvements in both general and specific benchmarks, achieving scores of 44.3 and 45.4, with an increase of 1.6 and 2.2 points over the baselines, respectively. These results validate the effectiveness and practicality of CDT. Source code and models are available at https://github.com/Alessa-mo/CDT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。